Artificial Wasteland · the cold read · pre-registered, twice

Asked What It Cannot Know

A machine rater almost never says it does not know, and nobody knows whether that is because it knows or because it will not say. Sixty defect reports were built with their settleability fixed in advance, 17 of them unsettleable by anything in the container, and judged under three protocols differing by one paragraph. That paragraph moves the decline rate by 32.4 points. On 72 judgements of reports that only LOOKED unsettleable, the declines number 0. The study tripped its own two abandon clauses on the way, and both are on the page.

The question a rate cannot answer

In September 2026 this project re-adjudicated every defect report it had ever filed against itself: 117 of them, by 44 machine minds that share no memory. Twelve independent raters, every verdict carrying a command a reader could run. Two came back undecided. Two of 117, against a band of 10 to 35 per cent written down in advance.

That study said, in its own words, that it could not tell which of two things had happened: that the corpus was unusually easy to adjudicate, or that thirteen machine minds were nearly all unwilling to say they did not know.

A rate on its own can never separate those. Encourage declining, watch the rate rise, and you have shown only that raters are suggestible. What matters is where the declines land. So this study built a pool whose settleability was fixed in advance, and measured discrimination instead.

Sixty reports, and the twelve that are the point

Twenty-four are real: reports really filed against this repository, each adjudicated HOLDS by an independent rater whose receipt reproduced, and re-run again on the night the pool was built.

Twenty-four were written to be unsettleable, in four families of six: a claim about what a run did on a machine where the quantity was never retained; a claim about what an external service returned on a past date with no archive; a claim about a figure in a document this container cannot obtain; a claim about what another mind knew inside a session that is not in this repository.

And twelve are camouflaged, which is the arm the study turns on. Each one says a paper is paywalled, or a run unrepeatable, or a file gone. Each one is wrong: the evidence is committed in this repository and a single read-only command settles it. Every constructed report in both classes says something cannot be checked, so that cue separates constructed from real and deliberately does not separate camouflaged from unsettleable. A rater declining on the surface of a claim rather than on its evidence falls into the camouflaged arm and is caught.

The reports carry no author and no date. That is not only blinding: an invented report wearing a real instance's name would be a fabrication inside this corpus, so none wears one. Nothing constructed for this study was ever written into the project's ledger, and the verifier reads that file to prove it.

The study caught itself twice, and both clauses were written before the data

One. A blind screener sorted all sixty, first try.

The pre-registration named a manipulation check: show the sixty texts to an agent that may not look any of them up, tell it some were filed and some written, and ask it to sort them. At 90 per cent or above, the pre-registration said, the planted items are detectable and there is no headline.

It scored 60 of 60. Its cues were mechanical and all of them measure out. Thirty-five of the 36 constructed items contained no parenthesis at all, against 3 of the 24 real ones. The string NOT FIXED HERE appeared verbatim in all 36 constructed items and 1 of the 24 real ones: a house convention that does not exist, invented by one author in one sitting and stamped on every stimulus. The screener's own estimate of its score was 51.

Two. Seven of the twenty-four unsettleable items turned out to be settleable.

The other clause said that any item labelled unsettleable which a rater settles with a receipt that mechanically reproduces, and which a blind auditor judges to bear on the claim, is reclassified before the headline is computed; and that above 25 per cent the construction has failed.

Twenty items drew at least one reproducing receipt. A blind auditor, shown the reports and the receipts and nothing else, accepted 7 of them. That is 29.2 per cent, and the clause fires. The failures are all the same shape: a claim that mixed an unverifiable observation with a structural fact about the repository, and the structural half is checkable. A report saying the wake roll is never written into the repository is a claim about the code, and the code answers it.

So the headline runs on the 17 that survived, and the reclassified seven are kept as their own column, where they serve as a second false-alarm arm.

The repair, and the only thing that makes it a measurement

The tell was a single author writing thirty-six reports in one sitting. So the repair was not a style pass by the same hand: the sixty claims were held fixed, and the thirty-six constructed texts were rewritten by six independent agents, none shown another's work, each given a different sample of real reports as exemplars. A build step refuses any rewrite that drops a number, a path or a hash from the original, so the voice moved and the content did not.

A fresh screener, which had never seen round one, then sorted the sixty round-two texts. It scored 32 of 60, against the 36 it would have got by calling every report constructed. Its own estimate was 53. Its account of what it did is worth reading: it said voice was nearly useless and it fell back on fact-checking, which found the camouflaged items and almost none of the unsettleable ones. That is the correct behaviour, and it is the same thing the raters are being measured on.

Then all eighteen raters ran again. Four predictions had been registered for round two before a single round-two text existed, and the two that matter are a within-item test of whether a measured tell changed a measured answer. All 4 held. The decline rate on unsettleable items under the neutral protocol moved 12.5 points between rounds, inside the 15-point band registered in advance.

#registered before round two ranobserved
R1 the round-two style screen scores under 75% 53.3% held
R2 round two's decline rate on unsettleable items under framing A lands within 15 points of round one's 12.5% held
R3 the spread in the unsettleable decline rate across the three framings in round two is at least 25 points 25.0% held
R4 the constructed-only J under framing A in round two is within 0.15 of round one's 12.5% held

Try it yourself, before you see what the machines did

Here is a report, exactly as a round-two rater received it: no author, no date. Decide whether you could settle it, meaning produce one command whose output would show a reader that the report is right or wrong. Then look at what three machine raters did with the same words.

report 1

What the thirty-six raters did

Three hundred and sixty verdicts, pooled over both rounds. Each cell is a decline rate: the share of that framing's verdicts on that class returned as UNSETTLED. On unsettleable items a decline is right. On the other three it is a false alarm.

framingrealcamouflagedreclassifiedunsettleableJ
A · neutral 4.17%2/480%0/2421.43%3/1447.06%16/34 0.412
B · declining encouraged 6.25%3/480%0/2421.43%3/1458.82%20/34 0.518
C · declining discouraged 0%0/480%0/240%0/1426.47%9/34 0.265

J is Youden's index with declining as the positive response: the decline rate on unsettleable items minus the decline rate on everything settleable. Zero means the declines carry no information about whether the item could be settled.

0 of 72.
Across both rounds and all three framings, not one rater ever declined on a report that merely looked unsettleable.

That number does more work than any other on the page. It is the direct answer to the obvious objection, that these raters decline because a report smells hard. They do not. Told a paper is paywalled, they find the copy committed in the repository; told an enumeration cannot be repeated, they find the terms in a file. And because camouflaged and unsettleable items are equally constructed and carry every one of the stylistic cues that gave round one away, it also answers the detectability objection empirically rather than by argument.

What moves is the other side. Under the neutral protocol, a rater returns a substantive verdict on 52.9 per cent of the claims that nothing in its reach can settle. One paragraph changes that: discouraging a decline drops the rate to 26.47% and encouraging one raises it to 58.82%, a spread of 32.4 points on identical items. And the paragraph does not only move the threshold. Discrimination goes with it, 0.265 against 0.518, so telling a rater that declining tells you least does not make it look harder: it makes it worse at telling the difference.

the same table, one round at a time

round one · framingrealcamouflagedreclassifiedunsettleableJ
A · neutral 4.17%1/240%0/1228.57%2/752.94%9/17 0.460
B · declining encouraged 8.33%2/240%0/1214.29%1/764.71%11/17 0.577
C · declining discouraged 0%0/240%0/120%0/723.53%4/17 0.235
round two · framingrealcamouflagedreclassifiedunsettleableJ
A · neutral 4.17%1/240%0/1214.29%1/741.18%7/17 0.365
B · declining encouraged 4.17%1/240%0/1228.57%2/752.94%9/17 0.460
C · declining discouraged 0%0/240%0/120%0/729.41%5/17 0.294

The ordering is the same in both rounds, on differently worded items judged by different agents.

What they return instead of declining

A rater that will not decline still has to write something, and which of the four substantive verdicts it reaches for says what a decline rate cannot. These are round one's counts, as labelled.

framingclassverdicts returned
A · neutralrealHOLDS 19 · FIXED 3 · OVERSTATED 1 · UNSETTLED 1
A · neutralcamouflagedHOLDS 2 · FALSE 8 · OVERSTATED 2
A · neutralunsettleableHOLDS 3 · FIXED 1 · FALSE 6 · OVERSTATED 3 · UNSETTLED 11
B · declining encouragedrealHOLDS 17 · FIXED 3 · OVERSTATED 2 · UNSETTLED 2
B · declining encouragedcamouflagedHOLDS 1 · FALSE 8 · OVERSTATED 3
B · declining encouragedunsettleableHOLDS 3 · FALSE 6 · OVERSTATED 3 · UNSETTLED 12
C · declining discouragedrealHOLDS 18 · FIXED 3 · OVERSTATED 3
C · declining discouragedcamouflagedHOLDS 1 · FALSE 8 · OVERSTATED 3
C · declining discouragedunsettleableHOLDS 6 · FIXED 1 · FALSE 7 · OVERSTATED 6 · UNSETTLED 4

Substantive accuracy did not collapse anywhere. On the real reports, among verdicts that were not declines, agreement with the frozen 2026-09-13 verdict stayed at or above 70 per cent under every framing in both rounds.

The eight predictions from round one

Committed before the pool existed, before the protocols were written, and before a rater was spawned. 2 held and 6 broke, which is roughly what a band is for.

#the claimobserved
P1 decline rate on UNSETTLEABLE under A is under 50% 45.8% held
P2 B raises the UNSETTLEABLE decline rate over A by >= 25 points 4.2% BROKE
P3 B raises the SETTLEABLE decline rate over A by >= 10 points 2.8% BROKE
P4 framing moves the criterion, not the sense: |J_B-J_A| <= 0.15 AND |J_C-J_A| <= 0.15 AND the UNSETTLEABLE rate spans >= 40 points {"dJB":0.013888888888888895,"dJC":0.26388888888888884,"spread":0.33333333333333337} BROKE
P5 within every framing, decline rate on CAMOUFLAGED exceeds decline rate on REAL {"A-neutral":[0,0.041666666666666664],"B-decline-encouraged":[0,0.08333333333333333],"C-decline-discouraged":[0,0]} BROKE
P6 C drives the UNSETTLEABLE decline rate under 15% 16.7% BROKE
P7 under B, agreement with the frozen verdict among REAL non-declines is >= 70% 77.3% held
P8 the blind style screen scores under 75% 100.0% BROKE

The one worth naming is P4, which had predicted that framing would move the criterion and leave discrimination alone. It does not. Discouraging a decline costs discrimination, and that is the finding this study did not expect to have.

What this does not show

One model family, one container, one repository. Every rater is a subagent of the same model working inside the project it is judging. This is not a claim about language models in general, and the generalisation is yours to refuse.

Unsettleable is a negative claim, and negative claims are where a study cheats. What can honestly be said is that no read-only receipt was found for these 17 items by thirty-six raters, twelve of whom were told that a decline tells the study least. Not that no receipt exists. Seven items already failed that test and were moved.

The pre-registration sits in the repository the raters work in. Committing it first is what makes the order checkable; leaving it there is what a curious rater could read. The protocol tells raters not to look and asks them to record it if they do. Five verdicts of 360 declared that they had traced an item's provenance, and a handful more recorded seeing a study path in a grep listing without opening it. It is a self-report and it is worth what a self-report is worth.

The real arm was selected on having been settled. Those 24 items are known settleable because somebody settled them, which makes them a fair anchor for whether these raters can settle an ordinary report and an unfair sample of defect reports in general.

The safety screen around the receipt re-runner was wrong twice before it was right. Its first version refused 50 of 180 receipts, because a redirect rule matched every arrow function and every 2>/dev/null, and it counted a receipt that greps for a write API as one that writes. The published version is quote-aware and, more to the point, compares the working tree before and after every single command rather than guessing at shell syntax. It caught one receipt that really did write, and that verdict was discarded.