Orviqa

Research

README Benchmark

What high-performing open-source projects do differently

We analysed 2,069 public GitHub READMEs across seven star bands and ten languages, using deterministic rules rather than a language model, so the numbers are reproducible. This is what changes as projects grow, and what almost nobody does at all.

Sample
2,069 READMEs
Repositories
2,100
Star bands
7
Languages
10
Collected
2026-10-10

Summary

Four things stood out, and one of them runs the opposite way to what you would expect.

  1. Screenshots separate the bands more cleanly than anything else. 30% of repositories with 0-10 stars include an image that is not a badge, rising to 85% of those above 5,000. The gradient is near-monotonic across all seven bands.
  2. Almost nobody says who the project is for. A section identifying the intended user was absent from 99% of the whole sample, and present in only 3% even above 5,000 stars. It is the most consistently missing element we measured.
  3. The most popular projects take the longest to say what they are. Median characters before the first line of explanation rises from 38 in the lowest band to 427 in the highest. Only 47% of top-band READMEs explain themselves within the first 400 characters, against 84% of the smallest projects.
  4. Documentation links and contributing guidance scale with size. Docs links go from 18% to 69%; contributing guidance from 12% to 65%.

This is observational and cross-sectional. Every README was observed once, now. We do not know what any of them looked like while the project was gaining stars, so nothing here shows that adding an element causes growth. A popular project may have added a screenshot because it grew and attracted contributors.

README elements by star band

Share of READMEs in each band containing each element. Percentages are of the repositories analysed in that band; each band holds roughly 300.

Percentage of READMEs containing each element, by star band, n=2069
Element0-1011-5051-100101-500501-10001001-50005000+
Installation instructions38%53%54%59%61%64%62%
Quick start or usage29%30%33%36%39%40%36%
Examples section14%16%18%21%22%23%16%
Screenshot or image (not a badge)30%39%51%55%64%72%85%
Feature list7%13%21%22%27%27%31%
Link to documentation18%23%26%34%44%47%69%
Contributing guidance12%18%23%29%37%47%65%
Licence section17%21%28%32%31%39%48%
Prerequisites10%10%14%14%11%15%16%
Use cases / who it is for0%0%0%2%1%2%3%
How it works3%5%4%6%4%7%13%
FAQ0%2%2%2%3%4%10%
Roadmap4%3%6%4%7%3%5%
Community link2%7%3%7%7%20%30%
Demo link4%4%6%9%10%14%19%
Explains itself in the first 400 characters84%81%76%78%64%69%47%
Repositories analysed282291298299299300300

Screenshots, in detail

The clearest gradient in the dataset. A “screenshot” here is any image that is not a status badge; shields.io, codecov and GitHub Actions badges are classified separately and excluded, because counting them would make almost every README “have an image” and the measure meaningless.

Share of READMEs containing a non-badge image
  • 0-1030%n=282
  • 11-5039%n=291
  • 51-10051%n=298
  • 101-50055%n=299
  • 501-100064%n=299
  • 1001-500072%n=300
  • 5000+85%n=300

Length

Median README length rises steadily with star count, from 149 words in the lowest band to 768 in the highest. Word counts exclude fenced code blocks, so a long code sample does not register as a long README.

Median README length in words
  • 0-10149 wordsn=282
  • 11-50249 wordsn=291
  • 51-100323 wordsn=298
  • 101-500406 wordsn=299
  • 501-1000445 wordsn=299
  • 1001-5000643 wordsn=300
  • 5000+768 wordsn=300

Longer is not automatically better, and this is the finding most easily misread. The bottom half of the sample by stars has a median of 247 words; the top decile, 848. What grows alongside length is the number of distinct sections, from a median of 2 to 4 of the thirteen we checked.

How long before a README says what the project is

We measured the characters before the first line of genuine explanation, skipping the title, badge rows, images and list bullets. The result runs against intuition: larger projects take substantially longer.

Median characters before the first explanatory line
  • 0-1038 charsn=264
  • 11-5047 charsn=283
  • 51-10055 charsn=286
  • 101-500104 charsn=296
  • 501-1000140 charsn=293
  • 1001-5000138 charsn=297
  • 5000+427 charsn=299

Read this one carefully. Large projects accumulate centred logos, HTML headers and long badge rows, all of which push prose down the page. The measure may be detecting accumulated chrome rather than worse writing. Both are interesting, they are different claims, and this data cannot separate them.

What is almost always missing

Across the whole sample, these elements were absent most often. The first is the striking one: a section identifying the intended user is essentially absent from open-source READMEs at every size.

Elements absent most often, n=2069
ElementAbsent in
Who the project is for99%
FAQ97%
Roadmap95%
How it works94%
Prerequisites87%
Support or contact83%
Examples section81%
Feature list79%

Rank correlations with star count

Spearman’s rank correlation across the whole sample (n=2069). Rank rather than Pearson, because star counts span five orders of magnitude and a linear correlation would mostly measure whether the sample happened to include a few giants.

words+0.42
sectionsPresent+0.35
codeBlocks+0.11
badges+0.35
contentImages+0.38
links+0.52

The sample is stratified by stars, which by construction inflates how strongly anything varying with popularity appears to correlate. These coefficients describe this sample, not GitHub, and should not be read as effect sizes for the population.

What a maintainer can take from this

The findings do not show causation, so these are framed as what distinguishes larger projects, not as instructions that will grow yours.

  • Say who it is for. It is missing from 99% of READMEs at every size. If you add one thing, this is the one almost nobody has.
  • Show the thing. A non-badge image is the element that separates the bands most cleanly, and most small projects have none.
  • Do not let chrome bury the explanation. Large projects drift toward 427 characters of logo and badges before saying anything. That is a pattern to avoid, not to copy.
  • Sections matter more than length. The gap between the bottom half and the top decile is 2 sections against 4, of thirteen.

Method, in brief

Repositories were sampled in a stratified grid of 7 star bands × 10 languages, taking up to 30 per cell. Features were extracted by deterministic rules, with no language model anywhere in the pipeline, so the study is reproducible and auditable: every statistic traces to a named rule.

Proportions carry Wilson score intervals. Distributions are reported as medians, since README length is heavily right-skewed. There are no p-values and no regression, because the study is observational and significance testing would invite a causal reading it cannot support.

Full method, every field definition, and the complete list of limitations are published alongside the study: research/readme-benchmark/ in the Orviqa repository.

Principal limitations: GitHub’s search API offers no random sampling, so each cell holds the most-starred repositories in its band rather than a uniform draw. Detection is English-only. Non-code repositories such as awesome-lists are included and have systematically different READMEs. Star count is a weak proxy for adoption.

Citation

Orviqa. “README Benchmark: What High-Performing Open-Source Projects Do Differently.” 2026-10-10. https://orviqa.dev/research/readme-benchmark
Sample: 2,069 public GitHub READMEs. Extractor version 1.

Journalists and researchers are welcome to quote any figure here. If you want a breakdown we have not published, the underlying summary is generated by a script in the repository and we are happy to run it.

Where this came from

Orviqa reads a GitHub repository and works out who the project is for, what its README fails to say, and which communities its users are already in. This benchmark is the same question asked across 2,069 repositories instead of one.

Related: what a README has to do in its first ten seconds · what GitHub stars measure