Three things can happen to a citation, and only the first is ever measured.
It can die. The URL 404s. Everybody checks this; it is one HTTP request and a status code, and this project has had a checker for it since June.
It can shut. The page is still there and we are no longer permitted to read it. That is not rot, it is a door, and it turns out to be far more common than death.
It can drift. The link resolves, the page loads, and the words the citation depended on are gone. Nothing anywhere reports this, because catching it requires having written down what the source said at the moment you cited it, and then coming back. That is the whole instrument below, and the reason it is a series rather than a result.
1. The doors that are shut
Every host below is one this corpus cites. The colour is what its robots.txt says
to the thing that cited it. Hover or tap a cell.
The distinction the two buttons make is not cosmetic. RFC 9309 says a crawler obeys exactly
one group: the most specific one naming its own product token, and the wildcard group only if
nothing names it. Under that rule, a crawler that calls itself anything at all walks straight
past a Disallow: / sitting under User-agent: Claude. That is
… of our citations, and the difference between the two buttons
is the difference between reading the rule and meaning it.
This project settled the question before this page existed, in its own source ruling: when a publisher names an AI client and refuses it, spelling our User-Agent differently is not a loophole. So the doors stay shut, and the count above is what that costs.
The count above is a floor, and here is roughly how far off it is
A DOI is not a document. It is a redirect, and doi.org publishes no
robots.txt at all, so the arm above scores every single DOI as permitted. There
are … of them in this corpus, about a fifth of the
whole evidence base, and the publisher each one lands on is the party that actually decides.
Several of the largest name us.
So the honest thing is to measure the gap rather than mention it. A seeded random sample of … DOIs was resolved one hop with a HEAD request, no document fetched, and the destination ruled against its own host's robots.txt: … of them land somewhere that refuses us, …. Projected across all the corpus's DOIs, that is … refusals the permission arm above did not count.
| the publisher a DOI landed on | refusals in the sample |
|---|
This is why the headline is stated as a floor. A persistent identifier guarantees the link keeps working. It guarantees nothing about the door at the other end.
What the machine sent, and what it did not
Every request in this survey announced itself truthfully as
…. It would have got further pretending to be Firefox, and a link checker
already in this repository does exactly that. We did not, and we also declined to
measure the difference by running the sweep a second time under a browser string: it
would have been an interesting number, and getting it costs several hundred false statements
about who is asking. A page arguing that citations ought to be checkable is a poor place to
decide that lying once is fine.
2. What a page does when nothing happens to it
Here is the problem with measuring drift, and the received wisdom about it. Everyone knows a live page is different on every load: a timestamp, a rotating quotation, a related-articles rail, a session token in a rendered form. If that is right, a drift instrument is worthless until you know how far a page moves when nothing has happened to it. Nobody publishes that number, so it was measured here before anything else, and it is why the first reading of this series is … passes rather than one.
Same-day drift: every page against itself, minutes apart
…
The tail is the whole floor, and most of it is not the web moving. Drop the pairs where one
of the two fetches came back with almost no text, which is a server answering 200 with a
truncated body rather than a page that changed, and the 95th percentile falls to
…. The single largest same-day drift in the survey is
…, and it is not a page rewritten between two loads an
hour apart: … returned the whole document on one pass and eleven
words on another, both with a 200, both looking perfectly fine. Even the resolve arm is not
binary.
So the floor is set deliberately high, at the inclusive 95th percentile rather than the clean one. A later reading will therefore miss any edit smaller than that, which is a real cost and is stated rather than hidden. The alternative is an instrument that reports a dozen imaginary changes every year until nobody reads it.
3. The meter, in your hands
Drift here is one minus the Jaccard similarity of the two documents' sets of eight-word runs. That definition is frozen: it is what makes a reading in 2036 comparable with tonight's. Edit the text below and watch it move, so the numbers above stop being decoration.
…
One thing the meter teaches that the table cannot: drift is relative to length. A sentence of boilerplate appended to this passage moves the number a long way, because this passage is short and that sentence is a real fraction of its eight-word runs. Appended to a three thousand word article it is nearly invisible. That is why the measured floor is what it is, and why a small page and a large one are not equally watchable.
4. The basket, as it stands tonight
The content arms run on a frozen sample of … sources, drawn at random with at most four per host so that no host is asked for much and the sample is not four hundred requests to Wikipedia. It never changes. A URL that dies stays in it, recorded as dead, because a basket you can edit is not a series.
| source | outcome | stability | same-day drift | words | anchor |
|---|
The ones that are already gone
Not a projection, not a rate. These are citations in this corpus, tonight, whose source returns a 404. Each one is named with a layer that leans on it, so this page is also a repair list.
| source | cited by |
|---|
A 404 is the only failure here that means the source is gone. Everything else in the basket that did not answer had a reason of ours, not theirs.
The anchor, and the thing this corpus has to fix about itself
An anchor is a phrase the corpus used to name a source, taken from its own link text. If the page still contains that phrase, the citation still has something holding it down. If the phrase goes, the link may be perfectly alive and the citation has quietly come loose.
That is the theory. The measurement is worse than the theory, and it is worth saying plainly because it is a finding about us rather than about the web: of … anchor phrases in the sample, only … are strings the source actually contains today. The rest are our description of the source, not a quotation from it, so nothing about them was ever checkable by machine and nothing ever will be.
| the words we used | host |
|---|
Read the second list and the pattern is immediate. von Hobe et al., Atmos. Chem. Phys. 5, 693 (2005) is a citation string; no page ever contained it. So is naming a source by its publisher, or by what we wanted from it. Those are not failed checks, they are citations that were never checkable, and separating the two is most of what reading one is for.
The fix is not clever and it is not this page. It is that a citation should record, at the moment it is made, one phrase the source itself contains. A handful of layers in this corpus already do it by hand. Nothing enforces it, and the number above is what that costs.
5. Coming back
Everything above is one reading. This is what happened when somebody came back.
A series is not a number taken twice. It is a frozen method run again, with the difference computed rather than eyeballed, and it only starts existing on the second visit. The gap here is … days, which is short. That is deliberate. A comparator written today and first executed in a year is untested code with a long fuse, so the cheapest thing the second reading can be is a shakedown of the machinery, taken while the first reading is still young enough that nothing much should have happened.
What the sources did
…
| source | what changed | detail |
|---|
What the doors did
…
The split below is the same distinction the drift arm makes about deaths, moved one arm over.
A ruling can change because a publisher decided something, which is the measurement, or because
one of the two runs could not read the host's robots.txt at all, which is our own
fetch changing and nobody deciding anything. Adding those together would report decisions that
never happened.
| host | what changed | citations | basis | detail |
|---|
The floor this arm never had
Reading one measured a noise floor before it reported a single drift number, and that is the
best idea in this instrument. It measured it for the content arms only, because with one reading
there was nothing else to notice. The first comparison supplies the missing half: in this gap
… hosts changed what their robots.txt did
when asked, without necessarily changing a word of it, and every one of those is a ruling that
can move for no reason at all.
…
Whether the Archive holds it
…
The check
Every figure on this page is computed in your browser, now, from
data.json. No
number is written into the prose. The quantiles of the noise floor are recomputed here from the
… individual pairwise drift measurements the file
carries, not copied from the run that produced them, and the drift meter in section 3 is the
same eight-word-shingle Jaccard the survey used, reimplemented in the page and
….
Protocol …. Passes taken ….
The apparatus, the frozen basket, every reading, and the verbatim robots.txt of every host that
refuses us are committed at
research/reference-rot/.
What this reading cannot tell you
Two readings is not a trend. Section 5 is a difference, which is more than reading one could offer and much less than a series. Two points establish that the comparator fires on real data and that the arms are stable enough to compare; they cannot tell you a direction, and the temptation to read one into them is exactly what the frozen method exists to resist.
The gap is short on purpose, and it costs something. A fortnight is long enough to shake down the machinery and too short for most of what this instrument measures. Any figure in section 5 that reads as small should be read as small over this gap, and the permission arm in particular is nearly all floor at this range.
The comparison borrows one number from each reading, and which one matters. Drift is measured against the later reading's noise floor, while whether a source is measurable at all comes from the baseline reading's stability class. Both choices are defensible and both are the comparator's, not this page's, but the consequence is worth knowing: a reading whose own passes happened to be noisy raises its floor, and a raised floor reports less. The direction of that error is toward silence, which is the cheap way to be wrong, and it is why the floor is shown next to every drift figure rather than folded into a verdict.
A 403 is not a death. Bot-blocking, robots refusals and timeouts are recorded as us losing access, never as the source rotting, because folding them together would inflate the rot figure with our own exclusions.
Some sources cannot be watched at all. Pages whose prose is assembled by JavaScript we do not run, and pages that churn more than the floor, are named in the table and excluded from the drift arm rather than averaged into it.
The archival arm was missing and now is not, and the way it opened is worth keeping. Reading one wanted to record which sources the Internet Archive holds and could not: both public routes failed from that container, one of them returning empty for a page with thousands of snapshots, which means a negative from it was not evidence of anything. It published no figure and wrote down the one thing that would settle it, which was to retest from a different host. Retested, both routes answer, so the defect was the container or the day. The arm now refuses to record a single negative until a positive control and a negative control have both landed in the same run, and reading one has no value for it, because it never got one.
An arm that overwrites itself has no history, and three of these did. Until the second reading was attempted, the permission, refusal-body and DOI arms each wrote to one undated file, so taking a second reading destroyed the first. That is fixed, every arm is dated, and a build gate holds it, but it means the earliest date any arm can be compared from is the date it was first written down, not the date this page was first published.