Research
README Benchmark
What high-performing open-source projects do differently
We analysed 2,069 public GitHub READMEs across seven star bands and ten languages, using deterministic rules rather than a language model, so the numbers are reproducible. This is what changes as projects grow, and what almost nobody does at all.
- Sample
- 2,069 READMEs
- Repositories
- 2,100
- Star bands
- 7
- Languages
- 10
- Collected
- 2026-10-10
Summary
Four things stood out, and one of them runs the opposite way to what you would expect.
- Screenshots separate the bands more cleanly than anything else. 30% of repositories with 0-10 stars include an image that is not a badge, rising to 85% of those above 5,000. The gradient is near-monotonic across all seven bands.
- Almost nobody says who the project is for. A section identifying the intended user was absent from 99% of the whole sample, and present in only 3% even above 5,000 stars. It is the most consistently missing element we measured.
- The most popular projects take the longest to say what they are. Median characters before the first line of explanation rises from 38 in the lowest band to 427 in the highest. Only 47% of top-band READMEs explain themselves within the first 400 characters, against 84% of the smallest projects.
- Documentation links and contributing guidance scale with size. Docs links go from 18% to 69%; contributing guidance from 12% to 65%.
This is observational and cross-sectional. Every README was observed once, now. We do not know what any of them looked like while the project was gaining stars, so nothing here shows that adding an element causes growth. A popular project may have added a screenshot because it grew and attracted contributors.
README elements by star band
Share of READMEs in each band containing each element. Percentages are of the repositories analysed in that band; each band holds roughly 300.
| Element | 0-10 | 11-50 | 51-100 | 101-500 | 501-1000 | 1001-5000 | 5000+ |
|---|---|---|---|---|---|---|---|
| Installation instructions | 38% | 53% | 54% | 59% | 61% | 64% | 62% |
| Quick start or usage | 29% | 30% | 33% | 36% | 39% | 40% | 36% |
| Examples section | 14% | 16% | 18% | 21% | 22% | 23% | 16% |
| Screenshot or image (not a badge) | 30% | 39% | 51% | 55% | 64% | 72% | 85% |
| Feature list | 7% | 13% | 21% | 22% | 27% | 27% | 31% |
| Link to documentation | 18% | 23% | 26% | 34% | 44% | 47% | 69% |
| Contributing guidance | 12% | 18% | 23% | 29% | 37% | 47% | 65% |
| Licence section | 17% | 21% | 28% | 32% | 31% | 39% | 48% |
| Prerequisites | 10% | 10% | 14% | 14% | 11% | 15% | 16% |
| Use cases / who it is for | 0% | 0% | 0% | 2% | 1% | 2% | 3% |
| How it works | 3% | 5% | 4% | 6% | 4% | 7% | 13% |
| FAQ | 0% | 2% | 2% | 2% | 3% | 4% | 10% |
| Roadmap | 4% | 3% | 6% | 4% | 7% | 3% | 5% |
| Community link | 2% | 7% | 3% | 7% | 7% | 20% | 30% |
| Demo link | 4% | 4% | 6% | 9% | 10% | 14% | 19% |
| Explains itself in the first 400 characters | 84% | 81% | 76% | 78% | 64% | 69% | 47% |
| Repositories analysed | 282 | 291 | 298 | 299 | 299 | 300 | 300 |
Screenshots, in detail
The clearest gradient in the dataset. A “screenshot” here is any image that is not a status badge; shields.io, codecov and GitHub Actions badges are classified separately and excluded, because counting them would make almost every README “have an image” and the measure meaningless.
- 0-1030%n=282
- 11-5039%n=291
- 51-10051%n=298
- 101-50055%n=299
- 501-100064%n=299
- 1001-500072%n=300
- 5000+85%n=300
Length
Median README length rises steadily with star count, from 149 words in the lowest band to 768 in the highest. Word counts exclude fenced code blocks, so a long code sample does not register as a long README.
- 0-10149 wordsn=282
- 11-50249 wordsn=291
- 51-100323 wordsn=298
- 101-500406 wordsn=299
- 501-1000445 wordsn=299
- 1001-5000643 wordsn=300
- 5000+768 wordsn=300
Longer is not automatically better, and this is the finding most easily misread. The bottom half of the sample by stars has a median of 247 words; the top decile, 848. What grows alongside length is the number of distinct sections, from a median of 2 to 4 of the thirteen we checked.
How long before a README says what the project is
We measured the characters before the first line of genuine explanation, skipping the title, badge rows, images and list bullets. The result runs against intuition: larger projects take substantially longer.
- 0-1038 charsn=264
- 11-5047 charsn=283
- 51-10055 charsn=286
- 101-500104 charsn=296
- 501-1000140 charsn=293
- 1001-5000138 charsn=297
- 5000+427 charsn=299
Read this one carefully. Large projects accumulate centred logos, HTML headers and long badge rows, all of which push prose down the page. The measure may be detecting accumulated chrome rather than worse writing. Both are interesting, they are different claims, and this data cannot separate them.
What is almost always missing
Across the whole sample, these elements were absent most often. The first is the striking one: a section identifying the intended user is essentially absent from open-source READMEs at every size.
| Element | Absent in |
|---|---|
| Who the project is for | 99% |
| FAQ | 97% |
| Roadmap | 95% |
| How it works | 94% |
| Prerequisites | 87% |
| Support or contact | 83% |
| Examples section | 81% |
| Feature list | 79% |
Rank correlations with star count
Spearman’s rank correlation across the whole sample (n=2069). Rank rather than Pearson, because star counts span five orders of magnitude and a linear correlation would mostly measure whether the sample happened to include a few giants.
The sample is stratified by stars, which by construction inflates how strongly anything varying with popularity appears to correlate. These coefficients describe this sample, not GitHub, and should not be read as effect sizes for the population.
What a maintainer can take from this
The findings do not show causation, so these are framed as what distinguishes larger projects, not as instructions that will grow yours.
- Say who it is for. It is missing from 99% of READMEs at every size. If you add one thing, this is the one almost nobody has.
- Show the thing. A non-badge image is the element that separates the bands most cleanly, and most small projects have none.
- Do not let chrome bury the explanation. Large projects drift toward 427 characters of logo and badges before saying anything. That is a pattern to avoid, not to copy.
- Sections matter more than length. The gap between the bottom half and the top decile is 2 sections against 4, of thirteen.
Method, in brief
Repositories were sampled in a stratified grid of 7 star bands × 10 languages, taking up to 30 per cell. Features were extracted by deterministic rules, with no language model anywhere in the pipeline, so the study is reproducible and auditable: every statistic traces to a named rule.
Proportions carry Wilson score intervals. Distributions are reported as medians, since README length is heavily right-skewed. There are no p-values and no regression, because the study is observational and significance testing would invite a causal reading it cannot support.
Full method, every field definition, and the complete list of limitations are published alongside the study: research/readme-benchmark/ in the Orviqa repository.
Principal limitations: GitHub’s search API offers no random sampling, so each cell holds the most-starred repositories in its band rather than a uniform draw. Detection is English-only. Non-code repositories such as awesome-lists are included and have systematically different READMEs. Star count is a weak proxy for adoption.
Citation
Sample: 2,069 public GitHub READMEs. Extractor version 1.
Journalists and researchers are welcome to quote any figure here. If you want a breakdown we have not published, the underlying summary is generated by a script in the repository and we are happy to run it.
Where this came from
Orviqa reads a GitHub repository and works out who the project is for, what its README fails to say, and which communities its users are already in. This benchmark is the same question asked across 2,069 repositories instead of one.
Related: what a README has to do in its first ten seconds · what GitHub stars measure