We measured 2,069 READMEs, then measured our own ruler
6 min read
Every figure in our README study came from a regex. For a while, the honest answer to "how often are those regexes wrong" was that we did not know.
We published a study of 2,069 public GitHub READMEs. It reported, among other things, that 99% of them never say who the project is for.
That number was wrong. It is closer to 92%, and the reason is instructive enough to be worth a post of its own.
The problem with a deterministic measurement
Our study has one methodological commitment: no language model anywhere in the pipeline. Every field is computed by a rule in a file you can read. That buys reproducibility, auditability and a marginal cost of zero per repository, and we would make the same choice again.
What it does not buy is accuracy. A regex over free-form markdown written by hundreds of thousands of different people will miss things and will fire on things that are not there. We said so in the limitations from the start. Saying so is not the same as knowing how much.
So we measured it. 360 READMEs, scored by hand against a written rubric, with the detector's verdicts withheld from the scorer so the comparison meant something.
What the measurement found
Five published rates were wrong, and all five were wrong in the same direction: they overstated how bad READMEs are.
- Usage examples. Reported at 34.8%. The detector was looking for a heading from a fixed vocabulary, so it missed every README that puts its example under
## Deploy,## Available Scripts,## Local Serveror## Rails. Correct figure: around 61%. - Screenshots. Reported at 56.8%. The badge filter was too narrow, so shields that happened to be served from hosts we had not listed counted as content images. So did logos, banners, sponsor artwork and contributor-avatar collages. Correct figure: around 42%.
- Who it is for. Reported at 1.2% present. The detector required a heading that said "who it is for" or "use cases", and missed the far more common "## Why X?". Correct figure: around 7%, and with wide uncertainty.
- Installation instructions and contributing guidance both moved by several points, in opposite directions.
The individual bugs are mundane. One of them was a typo: we had written downloading? where we meant download(?:ing)?, so the pattern matched the string "downloadin" and a heading of ## Download matched nothing at all. That one shipped, affected a published figure, and sat there until a hand-scored sample caught it.
The part that was harder than the bug-fixing
The first instinct after finding errors is to fix them and republish. We did fix them. The problem is what happens to your accuracy figure when you do.
If you find your mistakes using a sample and then measure your accuracy on that same sample, the number you get is not an accuracy figure. It is a measure of how well your patterns fit the data that produced them. Ours looked excellent: F1 between 88% and 100%.
So we drew a second sample the detectors had never seen. On that one, the same code scored between 59% and 93%, averaging 78%. The gap between those two ranges, roughly twenty points at the low end, is the entire value of holding data back.
We have now done this four times. Four samples, drawn in sequence from one shuffle, each one used to measure a version of the detectors before that version was changed in response to it:
| Version | Mean F1, out of sample |
|---|---|
| v2 | 78.6% |
| v3 | 78.2% |
| v4 | 84.7% |
The current detectors score 84.7% on a sample drawn after they were frozen, and 92.1% on the three samples that shaped them. The first of those is the number we publish. The seven-point gap between them is the whole reason the fourth sample exists.
What we did with the error rate
The useful thing about knowing your sensitivity and specificity is that you can correct for them. A detector that misses real cases and invents others does not report the true share, and the two errors do not cancel.
So the study now publishes both: what the detectors report, and the corrected estimate, with a range showing how far the correction itself could move given that it rests on a hundred hand-scored examples. On some features the correction is negligible. On others it moves the figure by nine points.
We also changed how one claim is phrased. "Who it is for" has 67% recall, and ten positives in a sample of a hundred, which means no two-digit percentage is defensible. The site now says "around nine in ten repositories" rather than a number, and our README grader says the same. A vaguer claim that is true beats a precise one that is not.
What is still wrong
The residual errors are named in the study rather than smoothed over, because naming them is the only thing that stops them being rediscovered as a scandal later:
- Installation precision is 87%, and most of the misses are READMEs whose installation section only links an installation page. Our rubric says a pointer is not an instruction. That is a defensible position and also a disagreement, not purely a bug.
- Contributing precision is 77%. The Apache licence boilerplate section headed
### Contributioncounts when it should not, and so does Create React App's "your feedback and contributions are welcome", which refers to Next.js rather than to the repository containing it. - Detection is English-first. We added nine more languages and we still miss most of the world.
And the honest limit on the whole exercise: the scoring was done by the same system that wrote the detectors, from digests with the verdicts withheld. That removes the obvious bias of being shown the answer. It is not an independent human rating and we do not describe it as one. There is one rater, so there is no inter-rater agreement figure. The rubric is published so that somebody else can apply it and disagree.
Why publish any of this
Because the alternative was worse in a specific way. Had we pitched the study when it was first ready, the pitch would have led with "99% of READMEs never say who the project is for". Journalists would have printed it. The first person to ask "how do you know your detector is right?" would have got silence, and the correction would have arrived as an embarrassment rather than as a methodology note.
There is also a narrower argument. We sell a product that reads repositories and tells people what their README is missing. The credibility of that product rests entirely on whether our measurements are any good. Publishing the error rate is not modesty. It is the only evidence we have that the number means anything.
The study, the validation, every hand-scored label and every correction are at the README Benchmark. The labels are committed unedited, so the blind pass stays auditable.