AGENTS.md Field Study

Concept synthesizing the analysis of 100 top GitHub repositories’ AGENTS.md files, revealing quantitative patterns in structure, content, and governance practices.

Key Findings

Structural Patterns

  • Median length: 1,198 words (barbell distribution: 37% >1,500 words, 10% <150 words)
  • Section convergence: Strong heading normalization across corpus
  • Top sections: testing (22%), commands (14%), project overview (13%), architecture (11%), development workflow (10%)

Content Distribution

Share of corpus word count:

  • Architecture & structure: 18.9%
  • Testing & validation: 17.2%
  • Commands & setup: 12.8%
  • Dedicated dos-and-don’ts: 11.3%
  • Git & PR workflow: 10.9%
  • Code style & conventions: 6.7%
  • Error handling: 4.7%
  • Docs & references: 3%
  • Security: 1.2%
  • Unclassified: 13.3%

Rule Density

  • 86% of repos contain explicit “don’t” rules
  • 784 explicit negative-rule bullets across sample
  • 90% use strong modality (must/always/never)
  • 56 of 99 non-empty files contain 3+ don’t-bullets

Language Adoption

Primary languages in sampled repos:

  • TypeScript: 25%
  • Python: 20%
  • Go: 13%
  • JavaScript: 12%
  • Rust: 10%
  • Only 27% of top 1,000 repos have any AGENTS.md file

Relationships

Implications

  1. Early adoption phase: Low AGENTS.md prevalence suggests most repos still lack explicit agent guidance
  2. Converging standards: Strong section normalization indicates emerging best practices
  3. Rule maturation: Larger projects show higher rule density and specificity
  4. Opportunity for standardization: Quantitative baseline enables comparison and improvement

Open Questions

  • How do these patterns evolve over time within individual repos?
  • What is the correlation between AGENTS.md maturity and project health metrics (contributor velocity, defect rates, etc.)?
  • Do specific industries or domains show distinct AGENTS.md patterns?

Sources

  • raw/prompts/articles/coldtea-agents-md-field-study.md — original field study capture