Ground-Truth · the corpus checks its own bibliography

The Semicolon in the Ampersand

This ground prints 1,373 references at the outside world and had never asked any of them whether they were true. So we asked, and the answer came back wrong the first time, in exactly the way the asking was supposed to catch.

Every number on this site is pinned by a program that re-derives it. That is the whole apparatus: some four hundred published recipes, each one re-running the arithmetic of the page it sits beside, so that a reader never has to take a figure on trust.

A citation is the one claim that apparatus cannot reach. It points out, at a record somebody else keeps, and nothing in this repository has ever read that record. The corpus has written “Farquhar, von Caemmerer & Berry (1980)” and never once asked Crossref whether that is who wrote it.

This page is the asking. Every DOI and every arXiv identifier the corpus prints, resolved against Crossref, DataCite and arXiv, frozen into a dated snapshot, and then judged field by field by raters who could not see which machine had flagged what.

citations checked, across pages
carry a bibliographic error, adjudicated blind
of them serious enough to stop you finding the work
of the automated auditor's alarms were real

And the first time we ran it, the answer was wrong. Round one put the error rate at 4.2 per cent and found something vivid: one page appearing to credit twenty-one of its twenty-six references to the paper’s last author instead of its first, thirteen of round one’s twenty-three errors in that layer alone. Two raters, working blind and independently, agreed on 338 of 347 items. They were both right about what they were shown, and what they were shown was not what the page says. A bare semicolon in the extractor’s list of separators was matching the one inside &, and cutting every reference list in half at its ampersand.

1. Divided by what

There is no such thing as “the citation error rate” until you say what you divided by. The medical literature that measures this in human journals halves its own headline on that choice alone: the same twenty-eight studies give 25.4 per cent on a denominator of references and 17.0 per cent on a denominator of quotations. So this page hands you the dial rather than a number.

Read the bands carefully; they are not like for like. The human-journal studies count volume, issue and page numbers as well, so their net is wider than this one and our figure is a floor under a narrower definition. They also checked references against the original article rather than against a registry record. And the model-generated rows measure something different again: a citation produced from memory and never checked by anybody, which is not what any reference here is.

2. What the raters were shown

Sixteen of the citations in the blind sample changed when the extractor was repaired. Every one of them had been read, in round one, by two raters who had no way to know that the text in front of them had been cut. Here is what moved. The left column is the evidence round one supplied; the right is the same reference as the page actually prints it.

3. The alarm and the error

Two automated auditors ran over the same citations: the fixed-window one this repository already had, and the unit-based one built to improve on it. They disagree, and neither of them is the truth. Pick a cell to read the citations inside it.

instrumentflagsprecisionrecallwhat that means

Precision and recall are weighted back to the stratum each rated citation was drawn from, because the sample deliberately over-drew flagged citations and a raw ratio would report a machine as finding most of the errors when what happened is that we mostly looked where it pointed. The counts are small; the interval on these is wide and the direction is not in doubt.

4. The bibliography, open

Every citation the corpus prints, beside what the registry says it is. Search a page, an author, a journal, an identifier.

5. What was actually wrong

Ten adjudicated errors, and four identifiers that resolve at no registry at all. The second list is the one that matters most, because an identifier that goes nowhere is the failure a reader cannot route around. All four are repaired, on the four layers named below, in the same change that published this page; each of those layers’ own checks still passes. The corpus measured here is the corpus as it stood before that, because recomputing the rate against the repaired pages would report a rate for a population that no longer contains what the rate counts.

6. What this does not measure

Whether a cited work actually supports the sentence it is attached to. The field calls that quotation accuracy, as against citation accuracy, and the two are routinely conflated and run at very different rates. A registry record cannot answer it: Crossref will tell you who wrote a paper and will not tell you whether it says what somebody claims it says. Everything on this page is the narrower question, and the harder half is untouched.

Volume, issue and page numbers. The human studies count those; this does not. Widening the net would raise the figure, not lower it.

Whether the raters are independent of the thing they judged. They are instances of the same model family that wrote the corpus, working to a brief written by another. The two rounds here are the best evidence available for what that is worth, and it cuts both ways: their agreement with each other was near-perfect while the evidence was wrong, which is exactly what correlated raters look like.

The citations with no identifier at all. Books, statutes, newspaper archives, standards, museum records. This audit can only see a DOI or an arXiv identifier, and a large part of what this corpus cites has neither.

7. The check

    The program is node research/the-corpus-cites/verify-the-corpus-cites.mjs. It recomputes every figure above from the committed snapshot and the committed rating files, and fails if any of them moves. It also re-fetches a hash-ordered sample of the registry records and reports how far the frozen snapshot has drifted from the live registries, because a measurement taken against a moving source has a shelf life and ought to say so.

    This page audits itself. Its own reference list is inside the corpus, so the next run of the census reads it like any other. The figures below were produced by running the instrument over this page before it shipped.

    Sources, and what was checked

      Registry records were fetched from the Crossref REST API, from doi.org content negotiation (which answers for DataCite registrants such as Zenodo, Dryad and OSF where Crossref correctly does not), and from the arXiv Atom API. The snapshot is dated on the page and committed beside the code that made it.