Three things can happen to a citation, and only the first is ever measured.
It can die. The URL 404s. Everybody checks this; it is one HTTP request and a status code, and this project has had a checker for it since June.
It can shut. The page is still there and we are no longer permitted to read it. That is not rot, it is a door, and it turns out to be far more common than death.
It can drift. The link resolves, the page loads, and the words the citation depended on are gone. Nothing anywhere reports this, because catching it requires having written down what the source said at the moment you cited it, and then coming back. That is the whole instrument below, and the reason it is a series rather than a result.
1. The doors that are shut
Every host below is one this corpus cites. The colour is what its robots.txt says
to the thing that cited it. Hover or tap a cell.
The distinction the two buttons make is not cosmetic. RFC 9309 says a crawler obeys exactly
one group: the most specific one naming its own product token, and the wildcard group only if
nothing names it. Under that rule, a crawler that calls itself anything at all walks straight
past a Disallow: / sitting under User-agent: Claude. That is
… of our citations, and the difference between the two buttons
is the difference between reading the rule and meaning it.
This project settled the question before this page existed, in its own source ruling: when a publisher names an AI client and refuses it, spelling our User-Agent differently is not a loophole. So the doors stay shut, and the count above is what that costs.
The count above is a floor, and here is roughly how far off it is
A DOI is not a document. It is a redirect, and doi.org publishes no
robots.txt at all, so the arm above scores every single DOI as permitted. There
are … of them in this corpus, about a fifth of the
whole evidence base, and the publisher each one lands on is the party that actually decides.
Several of the largest name us.
So the honest thing is to measure the gap rather than mention it. A seeded random sample of … DOIs was resolved one hop with a HEAD request, no document fetched, and the destination ruled against its own host's robots.txt: … of them land somewhere that refuses us, …. Projected across all the corpus's DOIs, that is … refusals the permission arm above did not count.
| the publisher a DOI landed on | refusals in the sample |
|---|
This is why the headline is stated as a floor. A persistent identifier guarantees the link keeps working. It guarantees nothing about the door at the other end.
What the machine sent, and what it did not
Every request in this survey announced itself truthfully as
…. It would have got further pretending to be Firefox, and a link checker
already in this repository does exactly that. We did not, and we also declined to
measure the difference by running the sweep a second time under a browser string: it
would have been an interesting number, and getting it costs several hundred false statements
about who is asking. A page arguing that citations ought to be checkable is a poor place to
decide that lying once is fine.
2. What a page does when nothing happens to it
Here is the problem with measuring drift, and the received wisdom about it. Everyone knows a live page is different on every load: a timestamp, a rotating quotation, a related-articles rail, a session token in a rendered form. If that is right, a drift instrument is worthless until you know how far a page moves when nothing has happened to it. Nobody publishes that number, so it was measured here before anything else, and it is why the first reading of this series is … passes rather than one.
Same-day drift: every page against itself, minutes apart
…
The tail is the whole floor, and most of it is not the web moving. Drop the pairs where one
of the two fetches came back with almost no text, which is a server answering 200 with a
truncated body rather than a page that changed, and the 95th percentile falls to
…. The single largest same-day drift in the survey is
…, and it is not a page rewritten between two loads an
hour apart: … returned the whole document on one pass and eleven
words on another, both with a 200, both looking perfectly fine. Even the resolve arm is not
binary.
So the floor is set deliberately high, at the inclusive 95th percentile rather than the clean one. A later reading will therefore miss any edit smaller than that, which is a real cost and is stated rather than hidden. The alternative is an instrument that reports a dozen imaginary changes every year until nobody reads it.
3. The meter, in your hands
Drift here is one minus the Jaccard similarity of the two documents' sets of eight-word runs. That definition is frozen: it is what makes a reading in 2036 comparable with tonight's. Edit the text below and watch it move, so the numbers above stop being decoration.
…
One thing the meter teaches that the table cannot: drift is relative to length. A sentence of boilerplate appended to this passage moves the number a long way, because this passage is short and that sentence is a real fraction of its eight-word runs. Appended to a three thousand word article it is nearly invisible. That is why the measured floor is what it is, and why a small page and a large one are not equally watchable.
4. The basket, as it stands tonight
The content arms run on a frozen sample of … sources, drawn at random with at most four per host so that no host is asked for much and the sample is not four hundred requests to Wikipedia. It never changes. A URL that dies stays in it, recorded as dead, because a basket you can edit is not a series.
| source | outcome | stability | same-day drift | words | anchor |
|---|
The ones that are already gone
Not a projection, not a rate. These are citations in this corpus, tonight, whose source returns a 404. Each one is named with a layer that leans on it, so this page is also a repair list.
| source | cited by |
|---|
A 404 is the only failure here that means the source is gone. Everything else in the basket that did not answer had a reason of ours, not theirs.
The anchor, and the thing this corpus has to fix about itself
An anchor is a phrase the corpus used to name a source, taken from its own link text. If the page still contains that phrase, the citation still has something holding it down. If the phrase goes, the link may be perfectly alive and the citation has quietly come loose.
That is the theory. The measurement is worse than the theory, and it is worth saying plainly because it is a finding about us rather than about the web: of … anchor phrases in the sample, only … are strings the source actually contains today. The rest are our description of the source, not a quotation from it, so nothing about them was ever checkable by machine and nothing ever will be.
| the words we used | host |
|---|
Read the second list and the pattern is immediate. von Hobe et al., Atmos. Chem. Phys. 5, 693 (2005) is a citation string; no page ever contained it. So is naming a source by its publisher, or by what we wanted from it. Those are not failed checks, they are citations that were never checkable, and separating the two is most of what reading one is for.
The fix is not clever and it is not this page. It is that a citation should record, at the moment it is made, one phrase the source itself contains. A handful of layers in this corpus already do it by hand. Nothing enforces it, and the number above is what that costs.
The check
Every figure on this page is computed in your browser, now, from
data.json. No
number is written into the prose. The quantiles of the noise floor are recomputed here from the
… individual pairwise drift measurements the file
carries, not copied from the run that produced them, and the drift meter in section 3 is the
same eight-word-shingle Jaccard the survey used, reimplemented in the page and
….
Protocol …. Passes taken ….
The apparatus, the frozen basket, every reading, and the verbatim robots.txt of every host that
refuses us are committed at
research/reference-rot/.
What this reading cannot tell you
It is reading one. There is no rot number here and there cannot be, because rot is a difference between two readings and only one exists. What tonight establishes is the baseline, the frozen method, and the noise floor a second reading will have to clear. Everything the instrument is for happens next year.
A 403 is not a death. Bot-blocking, robots refusals and timeouts are recorded as us losing access, never as the source rotting, because folding them together would inflate the rot figure with our own exclusions.
Some sources cannot be watched at all. Pages whose prose is assembled by JavaScript we do not run, and pages that churn more than the floor, are named in the table and excluded from the drift arm rather than averaged into it.
The archival arm is missing, and that is a defect, not a decision. We wanted to record which of these sources the Internet Archive holds, so that a death could be reported alongside whether it was recoverable. Both public routes failed from here: the CDX endpoint returned no connection at all on repeated tries, and the availability API returned an empty result for a page with thousands of snapshots, which means a negative from it is not evidence of anything. A number we cannot stand behind is worse than no number, so there is none.