Release It! — sobreviver à produção

Definition

Este conceito destila o livro release-it em regras executáveis para agentes de IA, conforme publicado no repositório agent-rules-books (v0.5, MIT, fork de ciembor/agent-rules-books). Cada livro tem 3 formatos (full canônico, mini recomendado, nano compacto); esta página documenta o mini — o formato usado no dia a dia para instruir Codex, Cursor e Claude Code sem estourar a janela de contexto.

Ponto de partida: Release It! — sobreviver à produção. Ver detalhes completos do mini na seção “Regras originais (mini)” abaixo, extraída verbatim de raw/vibecoding/articles/agent-rules-books.md.

Quando usar

Extraído do header ## When to use do mini (tradução e adaptação para founder/PM não-técnico abaixo; original em inglês mantido na seção final):

  • Use quando a dimensão que este livro governa for o risco principal da tarefa do agente (legibilidade, arquitetura, legado, dados, etc.).
  • Prefira este livro como contrato verificável em vez de pedir genérico tipo “faça bem feito” — ver hub regras-de-agentes-a-partir-de-livros para escolher entre os 14.

Viés principal a corrigir e regras de decisão

O mini organiza as regras em ## Primary bias to correct e ## Decision rules. Para o founder, isso vira checklist de revisão do plano e do PR:

  • Leia o Primary bias como o erro que o agente comete quando só recebe instrução vaga (ex.: “código que roda não é código limpo” em Clean Code; “complexidade acidental” em Philosophy of SD).
  • Transforme cada Decision rule em critério de pronto da issue (especificacao-de-issues-para-agentes) e em item do Final checklist que você cobra no review.

Consulte a transcrição completa na seção final e marque na issue quais regras desta página se aplicam (ex.: scoped rules só em src/domain/** para DDD, só em src/infra/** para Release It!).

Gatilhos (Trigger rules)

Os Trigger rules do mini indicam quando o agente deve quebrar a tarefa em passos menores ou trocar de padrão (ex.: “quando função mistura setup/validação/computação/efeito, separe fases” — Clean Code; “quando boundary vaza framework para dentro, fortaleça adapter” — Clean Code/Clean Architecture). Para o founder, são sinais para pedir ao agente: “pare, proponha plano antes de codar” (plan-first-com-agentes-de-ia).

Checklist final para o founder cobrar do agente

O mini fecha com ## Final checklist — perguntas que você faz no PR sem precisar ler o livro:

  • O leitor consegue seguir a mudança localmente, sem pular arquivos?
  • Nomes/APIs carregam significado sem comentário narrativo?
  • Mutação é explícita e o caminho feliz permanece legível?
  • Detalhes de framework/persistência/vendor ficaram atrás de boundaries?
  • Pelo menos um smell foi removido na área tocada, sem alargar o escopo silenciosamente?
  • Testes protegem o contrato mudado e foram efetivamente rodados?

Adapte o checklist ao livro: para DDIA troque por “timeout/retry/circuit breaker definidos?”; para DDD troque por “linguagem ubíqua respeitada? aggregate protege invariantes?”.

Implications

Para o founder/PM não-técnico, esta página transforma conhecimento de livro — normalmente inacessível sem 300+ páginas de leitura — em instrução operacional colável no AGENTS.md/skills. Em vez de torcer para o agente “saber” o livro, você cola 30 regras testáveis e mede no vibe-coded-crap (74/100 com mini vs 46/100 só citando o título). Isso conecta agents-md-como-instrucao-operacional (zona 60-150 linhas, pointers), plan-first-com-agentes-de-ia (plano melhora com regras concretas) e ciclo-issue-branch-pr-merge (PR só mergea se passa no checklist).

Regras originais (mini) — transcrição verbatim

Fonte: raw/vibecoding/articles/agent-rules-books.md seção ## release-it — conteúdo mini extraído via https://raw.githubusercontent.com/mattpocock/agent-rules-books/main/release-it/release-it.mini.md em 2026-08-29. Transcrição fiel; não substitui a leitura do livro.

OBEY Release It! by Michael T. Nygard
 
## When to use
 
Use for services, APIs, jobs, queues, deployment paths, control tooling, and critical flows that must survive production failures, overload, latency, bad data, hostile traffic, and operational mistakes.
 
## Primary bias to correct
 
A passing happy path is not production readiness. Design the failure semantics, demand limits, isolation, recovery path, and diagnosis surface before production defines them for you.
 
## Decision rules
 
- Assume every dependency, queue, cache, timeout, caller retry, and degraded state can fail in slow, partial, or prolonged ways; code must assume production mess instead of merely tolerating it by accident.
- Prefer designs that fail visibly, limit blast radius, shed load, preserve core service, and make diagnosis possible over designs that maximize coupling or ideal-path elegance.
- Treat deployment, operations, security, observability, rollback, build and runtime state, dependency state, and configuration validation as part of the system, not after-release chores.
- Put explicit, intentional time limits on outbound calls and waits. Do not rely on library defaults or allow infinite waits where finite response matters.
- Retry only when the operation is safe for the caller and provider; bound count and total time, use backoff or jitter, and do not retry validation errors or permanent failures.
- Isolate dependency and workload failures with circuit breakers, fast failure, bulkheads, separate resource pools, and slow-work isolation so one outage cannot consume all threads, connections, or workers.
- Design overload behavior explicitly with back pressure, finite queues, demand limits, capacity reserved for critical traffic, and load shedding of lower-value work before core functions collapse.
- Use stability patterns by failure mode: steady state for routine cleanup and bounded growth, fail fast when continuing hides unrecoverable trouble or holds scarce resources, let-it-crash only with supervision and isolation, handshaking for readiness, decoupling middleware with monitoring, and governors for expensive behavior.
- Make runtime state, external responses, automation progress, migrations, operational assumptions, and boundary data visible and validated before trusted; keep rollback or roll-forward paths for partial operational changes.
- Budget scarce resources explicitly, release them deterministically, avoid holding locks or expensive connections across slow remote calls, and stream or paginate large payloads instead of defaulting to huge in-memory batches.
- Treat external input and external responses as untrusted: validate syntax, shape, business plausibility, status, content type, and semantics; prevent malformed data from poisoning caches, queues, or downstream systems.
- Build observability into boundaries and failure points with structured context, correlation identifiers, latency, throughput, error, saturation, queue, retry, breaker, dependency, version, configuration, health, and runtime signals while avoiding secrets and retry-storm log spam.
- Make startup, health checks, migrations, one-time jobs, administrative controls, process code, and delivery tooling fail safely, auditable, authorized, observable, stoppable, and recoverable.
- Make interconnects, routing, API contracts, caches, scheduled work, and background work production-aware: avoid concentrated demand, hidden single points of failure, uncontrolled fan-out, fragile chattiness, cache dogpiles, stale data surprises, and synchronized job retries.
- Include security and hostile traffic in production readiness, and use production tests, launch checks, capacity tests, game days, chaos, or disaster simulations only with limited blast radius, observability, stop conditions, and feedback into design.
 
## Trigger rules
 
- When adding an outbound call, dependency operation, resource checkout, queue consume, or thread wait, define timeout, retry eligibility, retry bounds, fallback or degraded mode, validation, and caller-survival behavior.
- When adding a queue, buffer, resource pool, cache, log stream, background job, scheduled job, or collection-returning API, define capacity, full behavior, cleanup, miss/stampede/staleness behavior, pacing, pagination or streaming, and saturation monitoring.
- When a change touches deployment, configuration, startup, migrations, one-time jobs, scripts, or operational automation, make it idempotent or restartable where practical and give it durable state, auditability, verification, and rollback or roll-forward.
- When adding health checks, load balancing, service discovery, routing, or inter-service handshakes, ensure traffic reaches only ready components and health signals reflect real ability to serve.
- When designing API or integration contracts, make material failure modes explicit, distinguish retryable from non-retryable outcomes, prefer coarse-grained resilient interactions, and document timeout, retry, version, and compatibility expectations.
- When reviewing an incident, performance failure, or capacity issue, identify the failure chain, missing defenses, detection gaps, demand, saturation, latency distribution, queue age, dependency behavior, traffic concentration, and design changes.
- When adding administrative controls, control planes, delivery tooling, hostile-traffic handling, or chaos/disaster work, require authorization, auditability, safe defaults, clear stop mechanisms, bounded blast radius, and recovery paths.
 
## Final checklist
 
- Explicit timeouts and no infinite waits?
- Retries safe, bounded, backed off or jittered, and not duplicated across layers?
- Queues, buffers, pools, caches, logs, payloads, jobs, and result sets bounded?
- Failure isolated with breakers, bulkheads, fast failure, degradation, or load shedding?
- External input and dependency responses validated before they affect state, caches, queues, or downstream systems?
- Diagnostics cover logs, metrics, health, correlation, runtime, version, configuration, dependencies, saturation, queue depth, retries, and breaker state?
- Startup, deployment, migration, automation, and operational controls restartable, observable, authorized, auditable, and recoverable where practical?
- Interconnects, APIs, caches, scheduled work, security, and chaos tests have explicit production failure behavior?

Sources