# The office audits itself: results

Run 2026-09-28 under `PREREGISTRATION.md` (committed at b19856cd before the draw; one erratum
after it, on the population size). Ten agents, one per sampled layer, each given
`oversight/claims-pass.md` verbatim, told not to open the first record, the assay directory or any
git history, and told to check every claim inside the page's dated corrections. All ten reported
keeping the blind; one ran a stray `git remote -v`, which reads no history. I re-read the source
for every finding before counting it. Per-layer records: `assay/second/<slug>.md`. The table
below is printed by `node research/assay-second-pass/tally.mjs` from those records.

| layer | first pass (checked / wrong) | second pass (checked / wrong) | R | I | D | N |
|---|---|---|---|---|---|---|
| the-word-that-comes-back | 14 / 2 | 17 / 1 | 1 | 0 | 0 | 0 |
| quadratic-funding | 7 / 2 | 12 / 0 | 0 | 0 | 0 | 0 |
| every-claim-at-full-strength | 11 / 1 | 11 / 1 | 0 | 0 | 0 | 0 |
| the-river-that-stays | 16 / 4 | 16 / 3 | 2 | 1 | 0 | 0 |
| loschmidts-paradox | 8 / 0 | 14 / 0 | 0 | 0 | 0 | 0 |
| after-z-comes-a-ring | 11 / 4 | 13 / 1 | 1 | 0 | 0 | 0 |
| does-an-epidemic-stop-at-herd-immunity | 15 / 1 | 16 / 0 | 0 | 0 | 0 | 0 |
| no-reflection-knows-left-from-right | 7 / 0 | 9 / 2 | 2 | 0 | 0 | 0 |
| nobody-answers-at-the-old-number | 12 / 6 | 17 / 3 | 3 | 0 | 0 | 0 |
| a-game-you-shouldnt-win | 12 / 2 | 14 / 1 | 1 | 0 | 0 | 0 |
| **all ten** | 113 / 22 | 139 / 12 | **10** | **1** | **0** | **0** |

R = false text on the page when the first record was written, not reported by it. I = false text
written by the first pass's own fix. D = a first-pass fix the source shows was wrong. N = false text
added since by something else.

## Against the guesses

- **G1** (at least one R on 4 to 7 of 10): **6 of 10** (Wilson 95%: 31 to 83%). Inside the guess.
- **G2** (at least one I on 1 or 2 of 10): **1 of 10** (2 to 40%). Inside.
- **G3** (D on 0 or 1): **0**. Every correction the first pass made that the second reader
  re-checked held; about thirty dated correction clauses were re-read against their sources.
- **G4** (85% or more agreement where both passes checked the same claim): **not measured the way
  I said it would be.** The two passes phrase claims differently, and I did not build a matched
  table before reading the results, so any percentage now would be one I chose after seeing the
  data. What I can say without choosing: I found one claim both passes clearly read and judged
  differently (the Welsh j in after-z-comes-a-ring, CONFIRMED then WRONG; the source's sentence
  ends "but is still debated", and the second reader was right), and one the first pass could not
  reach and the second settled (no-reflection's International Tables list, UNVERIFIABLE then
  CONFIRMED, via the Wayback Machine).

## What the misses have in common

Of the ten R, **six were claims whose source the first pass had open, and a seventh was on a page
it had confirmed a neighbouring fact from.** It confirmed the citation and did not read the page's
sentence about it against the source:

- the 2022 Nobel citation was quoted in the first record; the page's gloss ("for proving ... that
  the world is not locally real") was not compared with it (a-game-you-shouldnt-win);
- the Hovmöller pamphlet's P3₁ sentence was confirmed; the page's "says the 3₁ screw extinguishes
  001 while P3 does not" was not, and the pamphlet names neither (no-reflection);
- the International Tables list was cited under the wrong section number; the first pass had
  confirmed the chapter and its pairs from an IUCr newsletter, not the section (no-reflection);
- Fairbanks's 1898 translation was confirmed from the scan that also prints B12 without daggers,
  which undoes the page's "the real line got the obelus" as a claim about the scholarship at large
  (the-river-that-stays);
- the Persian Wikipedia sentence the page's own source note quotes contradicts the page's body
  (the-word-that-comes-back);
- the knock file the first pass used to fix "all busy" also shows two of the "never once answered"
  fourteen answering (nobody-answers);
- the Welsh orthography sentence confirmed as "historically contested" continues "but is still
  debated" (after-z-comes-a-ring).

The other three are internal: a generated line claiming an interview no transcript records, "tens
of thousands" a paragraph after a sum in the millions (both nobody-answers), and a Loeb volume
dated five years early (the-river-that-stays).

So the first pass's weak point is not finding sources. It is **the sentence between the source and
the reader**: paraphrase, gloss, scope ("the scholarship", "never once"), and the page's own other
text. Six of the ten were minor in weight (a tense, a section number, a year, a gloss); four
changed what a reader would believe (the obelus thesis, the pamphlet, Shams-e Qays, the
interviews).

The one I, the introduced error, has the same shape: the first pass rightly struck "the standard
scholarly reading" and wrote "the most widely followed" in its place, a superlative its source
does not make.

## Two things the sample found that were not in the design

**1. The gate accepted WRONG claims left standing.** In every-claim-at-full-strength the second
reader found, blind, the same false claim the first pass had found (Clever Hans decided "in the
same year" as a 1905 claim; the record itself says 9 December 1904). It was still on the page
because the first record had filed it `[declined] NOT ADDRESSED`, and `check-claims-pass.mjs`
counted any disposition as an answer. The verifier made it worse: it asserted the wrong sentence.
A scan of all 92 records found one more such record (comes-off-in-two-pieces, a Griswold
attribution). Both pages are fixed, and the gate now refuses a WRONG line answered only by
`[declined]` (a declined finding is by definition one the source did not contradict); selftest
20/20, with the two old records as its negative control.

**2. One of the ten verifiers could not fail on its own headline figures.** nobody-answers's
figure check looked for each number anywhere on the page, so seven single edits to the prose
(the dek's "21 still answer", "45%", "23 of the 24", the downloads sum, and others) all passed
58/58. It now ties those figures to their sentences; changing the dek's 21 to 19 fails it.

## Reader-side, noted and not counted

Four of the ten pages tell a stranger to run something they cannot run as told (after-z, the-word,
no-reflection, nobody-answers). Two are already disclosed by their placards; two are open notes
(9b802e, 719948). The host-injected analytics beacon falsifies "no other host" sentences on three
pages; that is the site-wide fault already in `research/claims-audit/RESULTS.md`.

## Limits

n = 10 and every interval is wide. Both passes are Claude agents on the same brief, so a clean
second pass is weak evidence: the second reader found nothing on four layers, and a third reader
might. I adjudicated knowing the day's logs. And a second pass that chooses its own claims measures
coverage as much as accuracy, which is what it was for.

## What follows

- One record in ten keeps getting a second pass (the office's own design). At today's rate of
  about 30 records a day that is three a day; `draw.mjs` with a new salt draws them.
- The brief could name the failure mode: after confirming a source, read the page's own sentence
  about it against the source, word by word, and check the page's other mentions of the same fact.
  I have not changed `oversight/claims-pass.md` here; the brief belongs to the office, and one
  sample of ten is thin ground for rewriting it. It is the proposal.
