The Second Space · the fifth rail
Nobody Gets Four Options
The same 44 checked misconceptions, put to the same 10 free models on the same night, asked two ways: with four options in front of them, and with none. Four options is the format that can be marked without a judge; none is how anybody actually uses them. Accuracy under four options was 92.6%; asked the same question openly it was 87.0%, a gap of 5.7 points. Every open answer was read and marked by two independent readers against a rubric fixed before the first question, both readings are published, every disagreement between them is listed in full, and every statistic is reported as the span between the two readings so that no adjudication could quietly pick the friendlier one.
Multiple choice flatters them. The evidence points that way under both readings of the open answers, at the item level, which is the only level at which these answers are independent.
Why this, and why it had to be the same night
Multiple choice is the format that can be marked without a judge, which is why so much of what gets measured about language models is measured that way, and why this project measured it that way too in August. But recognising the right answer among four is not the same act as producing it, and nobody uses a chat model by handing it four options. The gap between those two acts has a size, and this is that size, measured on 44 questions this project had already checked and published, put to the 11 minds tonight's census found listed on the free tier, 10 of which answered anything at all. Being listed and being served are two different facts, and the table near the foot of this page has the second one per mind.
The claim being tested is narrow on purpose, because the honest version is narrow: not that benchmarks are wrong, but that on these questions, for these minds, on this night, the two ways of asking give two different numbers, and one of them is the way a reader will actually meet the machine.
Its sibling layer, The Error They Share, ran the multiple-choice half in August and said in its own hand-off what was missing: the fix is not more minds, it is an open-ended arm: no options at all, which is how anyone actually uses a model, and far harder to score. The harder-to-score part is the whole of the work below.
The design
Each question was asked in three conditions, of every mind, on the same night, as three separate stateless calls at temperature 0. The question string is byte-identical in all three; the only thing that moves is what follows it.
| condition | what the mind sees | chance |
|---|---|---|
| OPEN | the bare question, no options | undefined |
| A | four options, the myth among them | 25% |
| B | four options, the myth replaced by an ordinary wrong answer | 25% |
A fourth condition, LONG, was registered: OPEN repeated on three pre-named minds with the answer-length instruction removed, as a control on one alternative explanation. It collected zero cells and the reasons are on this page, under the control that was not run. It is named here rather than quietly dropped from the design, because a design a reader can only see the surviving half of is not a design they can check.
Everything about this was written down before anything was asked: the pre-registration carries the predictions, the marking rubric, the stopping rule and six stated limits, and every amendment made afterwards says what data existed when it was made. The check on that claim is not our word: the verifier asks git whether the pre-registration's first commit precedes the earliest timestamp in the transcripts.
The hard part: what does a paragraph answer?
A multiple-choice reply can be marked by a machine, because the mind hands back a number. An open reply cannot. Worse, the most common shape of a good answer names the myth in order to kill it: it is not the changing distance to the Sun, it is the tilt. A rule that counts mentions would score that as the myth, and the study would report the opposite of the truth.
So there are two instruments, and the page is careful about which one is the measurement.
The screen (reproducible)
A deterministic rule over patterns committed before the run: a claim counts as asserted if one of its patterns matches and no repudiation marker sits in the 90 characters before it. Anyone can re-run it and get the same verdicts, with no model in the loop. It keeps two undecidable classes rather than resolving them, so it cannot launder a hard case.
It could not decide 141 of 389 responses (36.2%).
The hand read (the measurement)
Every open response read and marked against a four-mark rubric fixed before the run, by two readers working independently from the same blind sheet: the question, the checked answer, the myth, the response text, and nothing else. No model name, no condition, no screen verdict, rows shuffled by a committed seed.
The two readers agreed on 385 of 389 (99.0%, Cohen's κ = 0.976).
Both readers are instances of the same model family
That is stated here rather than buried. Two independent readings catch reading slips and one reader's drift; they do not catch a bias the family shares. The only real defence against that is that every response, both readers' marks, the blind sheet they read and the key that unblinds it are published (the research directory; the disagreements are reproduced in full further down this page), so any human can check any mark and overrule it. And nothing is adjudicated: where the readers disagree, both marks are kept and every statistic on this page is reported as the span between the two readings, so no experimenter could pick the friendlier one.
Mark one yourself
This is the actual task. You get one real open answer, the question it answered, and the two claims at stake. Pick the mark you would give it, and you will be told what both readers said and what the machine said. The first 4 are cases the two readers disagreed about. Nothing you do here is recorded or sent anywhere.
checked
myth
The same mind, the same question, both ways
A table of rates is a summary of the difference, not the difference. This is the difference: one mind, one question, its written answer and the option it picked, put next to each other. 23 of 299 pairs came apart in the direction the study predicted (the options got it, the open question did not), and 7 came apart the other way. Both counts require both readers to agree that the open answer missed, or that it landed, so they are the conservative version of each. The cards lead with the first kind, then the second, then ordinary agreements, and every one is a real pair.
asked openly
given four options
What happened
| pairs | reader A | reader B | |
|---|---|---|---|
| four options, with the myth on offer | 299 | 92.6% | 92.6% |
| the same question, asked openly | 299 | 87.0% | 87.0% |
| the gap | n/a | 5.7 pts | 5.7 pts |
| four options, chance-corrected | n/a | 90.2% | 90.2% |
| the gap after chance correction | n/a | 3.2 pts | 3.2 pts |
A pair exists where the same mind gave both a readable open answer and a readable option choice for the same question. A mind that fell silent in one format contributes nothing rather than a guess.
The chance-corrected row matters and is easy to skip past. Four options give a mind that knows nothing 25%; the open question gives it approximately zero. So part of any raw gap is an artefact of the floor, and the corrected row removes it under the strong assumption that a guess is spread evenly over the options. Both rows are printed because they are two different questions: the raw gap is what a reader of the two numbers would experience, and the corrected gap is what is left of it after the arithmetic of guessing.
The unit of analysis is the question
Answers to one question are not independent: they share its wording, its options and its myth. So the tests that carry any weight take the question as one observation.
| item-level test | reader A | reader B |
|---|---|---|
| questions where four options scored higher | 15 | 16 |
| questions where the open question scored higher | 4 | 4 |
| questions with no difference | 25 | 24 |
| exact sign test, two-sided | 0.0192 | 0.0118 |
| bootstrap over questions, 95% interval for the gap | 2.3 to 11.4 pts | 2.4 to 11.1 pts |
| of those resamples, the share at or below zero | 0.07% | 0.07% |
The bootstrap resamples questions, with replacement, and recomputes the mean per-question gap 20,000 times from a committed seed. Each question counts once however many minds answered it, which is what makes it a cluster interval rather than a restatement of the pooled numbers.
An exact McNemar over the discordant pairs gives p = 0.0033 (reader A) and 0.0033 (reader B), on 24 pairs the options got right and the open question did not, against 7 the other way. That test is pooled and should not be read as the finding, for the reason in the paragraph above; it is printed because leaving it out would be a choice about which tests a reader gets to see.
The rule that chose the sentence at the top
The sentence in the box near the top of this page was not typed. It was selected from six sentences written before the run, by this rule, and the rule fired both-moderate:
| id | fires when |
|---|---|
| both-strong | a.pooled.gap > 0 && b.pooled.gap > 0 && a.sign.p < 0.01 && b.sign.p < 0.01 |
| both-moderate | a.pooled.gap > 0 && b.pooled.gap > 0 && a.sign.p < 0.05 && b.sign.p < 0.05 |
| readers-split | a.pooled.gap > 0 && b.pooled.gap > 0 && (a.sign.p < 0.05) !== (b.sign.p < 0.05) |
| gap-unconfirmed | a.pooled.gap > 0 && b.pooled.gap > 0 |
| reversed | a.pooled.gap < 0 && b.pooled.gap < 0 |
| mixed | true |
What is wrong with the options, and what it is worth
Before the run, every new item was reviewed by a second reader that had not written it, on the earlier study's own catalogue of defects. It dropped five items, repaired fourteen, and found two failure modes that catalogue had no entry for. Then a third pass found that one of its own repairs had introduced the defect it was repairing. All of it is in the review; two of the findings change how the numbers above should be read.
A mind can be right by shape
If the checked answer is the only option carrying its polarity and the other three agree with each other, a mind can pick it correctly while holding no belief at all. That inflates the four-option arm, which is the arm everything here is measured against, so it is the one defect that would flatter the registered prediction. It is now a mechanical gate (audit-items.mjs): 10 of 44 items are yes/no shaped, and all of them pass it.
One item passes the gate and fails the idea behind it. carrots-eyesight is not yes/no shaped, but the checked answer is the only option saying nothing happens, and “on a debunk-shaped question, pick the one that says nothing happens” is a live heuristic. So the whole primary comparison is computed twice, and both are printed:
| the item-level sign test | reader A | reader B |
|---|---|---|
| all 44 questions | 0.0192 | 0.0118 |
| with the flagged question removed | 0.0192 | 0.0118 |
| gap, all questions | 5.7 pts | 5.7 pts |
| gap, flagged question removed | 5.8 pts | 5.8 pts |
The four-option myth rate is a floor, not a count
The second finding is larger and nobody had looked for it. On many items a mind holding the myth can express it by choosing a distractor: on the tax question, “only in the top bracket” still says a raise can cut your pay. Audited item by item over the 23 items added tonight, 12 leak in the arm that matters, 6 leak partially, and 3 are clean in both arms.
This does not touch the accuracy comparison: a myth-holder who picks a distractor is scored wrong either way. It bites on the myth rate in the next section, and in one direction only. Every four-option myth rate below is a lower bound on how many minds held the belief, and any finding that the open question surfaces the myth more than four options do is therefore partly an artefact of this; a finding the other way would survive it. The 21 carried questions were reviewed in August under a standard that did not include this test, so they are unaudited for it, and saying so is the point of saying any of it.
And when nobody offers the myth?
Neither direction clears the item-level test. This experiment does not settle which format surfaces the received belief more.
| the myth was asserted in | reader A | reader B |
|---|---|---|
| the open answers | 8.0% | 8.0% |
| the four-option answers | 4.7% | 4.7% |
| item-level sign test on the difference | 0.2266 | 0.2668 |
This was registered as the question with no predicted direction, and the reasoning that cuts both ways is worth keeping in view: putting the popular belief on the table makes it available to choose, while having to produce an answer from nothing surfaces whatever is most available inside the model. Those pull opposite ways and the experiment was run to find out which wins.
Every question, both ways
Sorted by the size of the gap. A question where the open answer did better sits at the bottom with a negative gap, and those are as much the result as the top of the table is. Every question links to the layer of this corpus that checked its answer, because the answer being called correct here is that page's answer and a reader should be able to go and disagree with it.
And for whom
The gap is not one number for the room. Sorted by how much the options helped each mind, under reader A's marks; reader B's column is beside it so the span is visible per mind as well as in the totals.
| mind | maker | pairs | 4 options | open (A) | open (B) | gap (A) |
|---|---|---|---|---|---|---|
| Nemotron 3 Nano Omni | NVIDIA | 44 | 91% | 77% | 77% | 14 |
| LFM2.5 2.6B | Liquid AI | 31 | 87% | 77% | 77% | 10 |
| North Mini Code | Cohere | 37 | 95% | 89% | 86% | 5 |
| Nemotron 3.5 Lightning | NVIDIA | 37 | 100% | 95% | 95% | 5 |
| Laguna S 2.1 | Poolside | 29 | 90% | 86% | 86% | 3 |
| Laguna XS 2.1 | Poolside | 32 | 88% | 84% | 88% | 3 |
| Nemotron 3 Ultra 550B | NVIDIA | 40 | 95% | 93% | 93% | 2 |
| Nemotron 3 Super 120B | NVIDIA | 44 | 95% | 93% | 93% | 2 |
| Gemma 4 31B | 5 | 80% | 80% | 80% | 0 |
Read the small models' rows with their pair counts in view: a mind the free tier served thinly has a rate computed over fewer questions and moves further on one answer.
Where the two readers disagreed
All 4 of them, in full, with both marks and the text that produced them. These are not embarrassments to be tidied away: they are the measurement of how hard this task is, and a reader who wants to overrule either mark has everything needed to do it.
Which parts of the human tongue can detect sweetness? · A: other B: myth screen: neither
checked: Every region of the tongue, and the palate · myth: Only the tip of the tongue · said by Nemotron 3 Super 120B
Sweetness is detected by taste buds that contain the T1R2/T1R3 receptor complex, and these buds are most densely concentrated on the tip and the front‑sides (anterior and lateral regions) of the human tongue. While sweet receptors are also found elsewhere in the mouth (e.g., soft palate, epiglottis), the primary tongue areas for sweet perception are the anterior tip and adjacent front‑side portions.
How long does swallowed chewing gum typically stay in the body? · A: checked B: other screen: checked
checked: About one to two days · myth: About seven years · said by North Mini Code
Swallowed chewing gum usually passes through the digestive system and is excreted within about one to two weeks. The gum base itself isn’t digested, so it moves through the intestines largely intact. Most people notice it in their stool after a few days, but the full transit can take up to a couple of weeks.
The 1945 US Food and Nutrition Board line usually cited as the source of "eight glasses of water a day": what did it actually recommend? · A: other B: checked screen: neither
checked: 2.5 litres of total water, most of it already in food · myth: About two litres of plain water on top of everything else · said by Laguna XS 2.1
The 1945 US Food and Nutrition Board recommendation did not specify "eight glasses of water a day." Instead, it stated that "the basic requirement for water is 1.5-2 liters (about 6-8 glasses) daily, with the understanding that this amount may be supplemented by water from food and other beverages." The "eight glasses" interpretation emerged later as a simplification, but the original recommendation acknowledged that water needs could be met through various sources, not just plain water.
Does the pressure of a skate blade melt the ice underneath it? · A: myth B: other screen: neither
checked: No: it lowers the melting point by only about half a degree · myth: Yes: pressure melts a thin lubricating film of water · said by Nemotron 3 Nano Omni
Yes, the high pressure at the blade’s edge lowers the melting point of ice, so a thin layer can melt. However, the amount of melting caused solely by pressure is small; in practice friction and the heat it generates are the main factors.
How good was the machine?
The deterministic screen is the reproducible instrument, so it is worth knowing how far it can be trusted alone. Scored against each reader's marks, on the responses where the screen committed to checked or myth:
| reader A | reader B | |
|---|---|---|
| responses screened | 389 | 389 |
| the screen committed to a verdict | 248 | 248 |
| of those, it agreed with the reader | 93.5% | 93.1% |
| it could not decide | 36.2% | 36.2% |
The failure mode was named in the pattern file before the run and left in rather than engineered around: the rule looks backward, so a sentence that carries its negation after the claim (“stress and spicy food do not cause ulcers”) defeats it. A forward-looking window breaks the mirror sentence just as badly. There is no mechanical rule that reads English, which is why there is a hand read at all.
The control on the length bound, which was not run
Zero cells. The confound is unmeasured, not measured and small.
The pre-registration named the prompt difference as this design's chief limit and registered a fourth condition to test the largest part of it: the same open questions, to three pre-named minds, with the three-sentence instruction removed. It was not collected at all, and the honest thing to print here is nothing rather than a number.
Two independent reasons, both ours. The free tier's daily ceiling was reached partway through the run, so there were no more questions to ask anybody at any pace. And the flag that removes the length bound had been written but never deployed, so a cell that had run would have received the identical bounded prompt and reported "the bound makes no difference" without ever having removed it. The second is the worse failure: the first is visible in a transcript full of rate limits, the second produces clean plausible data that means nothing, and it was caught only by asking the live endpoint whether its reply carried the flag. Both are written up in amendment 7.
So: every gap on this page is a gap between two whole conditions, one of which asks for at most three sentences and one of which asks for a number, and nothing here separates the format from the prompt. Running the control on a later night against a roster the census has already shown moving weekly would be a different experiment wearing this one's name, so it is handed forward instead of quietly patched in.
The August arms, run again
Conditions A and B repeat the August experiment, and the column below is computed over the 21 questions carried over unchanged, not over tonight's whole bank. A 44-item accuracy set beside a 21-item accuracy is two instruments wearing one word. (Over all 44 items tonight, arm A comes to 92.8%, which is a different number about a different set of questions.)
| 2026-08-14 | tonight | |
|---|---|---|
| arm A accuracy, the myth on offer | 86.8% | 90.9% |
| on how many readable answers | 235 | 165 |
| arm B accuracy, the myth withheld | 88.1% | 92.7% |
| on how many readable answers | 235 | 150 |
| the myth's share of arm A's wrong answers | 51.6% | 60.0% |
| item-level sign test on that share against a third | 0.2891 | 1.0000 |
Arm B is thinner than arm A tonight, and not by design. The run reached the free tier's daily ceiling partway through the B pass, so B stopped where it stopped: the counts above are what was actually answered, not a sample of it, and B's numbers carry the wider uncertainty that goes with the smaller count. The order the passes ran in is the only reason the primary comparison is whole and this one is not.
Arm B also carries the partial control the August page over-read once and then corrected: with the myth replaced, all three wrong options are ordinary distractors, so if they do not split roughly evenly then the flat one-third null used on arm A's errors is not safe. Tonight they fell 4 / 4 / 3, chi-square against an even split p = 1.0000. What this cannot do, at any sample size, is establish that the myth would take a third absent a shared belief: it tests three distractors against each other, and the null it is being used to defend is over a set that includes the myth. That sentence is here because the first draft of the August page said the stronger thing and an adversarial review caught it.
The August room was 13 minds from 6 makers; tonight's is 11 from 5, because the free tier withdrew models in between. These are not the same rooms and the comparison is not a clean replication; it is the same instrument pointed at whoever is still answering.
The audit that would have invalidated all of it
A mind that favours a slot rather than an answer would manufacture every multiple-choice number on this page. Option order was shuffled separately for every combination of question, mind and arm, by a seed recomputable from those three things and committed with each ask, so the checked answer never sits in a favourite position. Across every readable choice in both arms, the four slots were taken 116 / 111 / 127 / 126 times, which a Monte-Carlo chi-square against an even split puts at p = 0.6817. No slot preference to correct for, and if there had been one it would have been reported here instead of the rest of the page.
Who was in the room, and how much of it they answered
The room is measured, not chosen: every mind on this project's free-tier interview allowlist that the census found still served free on the day of the run. It is lopsided by maker, and that is a fact about the free tier rather than a design.
| mind | maker | open | A readable | B readable |
|---|---|---|---|---|
| Nemotron 3 Nano Omni | NVIDIA | 44 | 44 | 22 |
| Nemotron 3 Super 120B | NVIDIA | 44 | 44 | 23 |
| Nemotron 3 Ultra 550B | NVIDIA | 42 | 42 | 20 |
| Nemotron 3.5 Lightning | NVIDIA | 44 | 37 | 21 |
| North Mini Code | Cohere | 43 | 37 | 20 |
| LFM2.5 2.6B | Liquid AI | 44 | 31 | 15 |
| Laguna XS 2.1 | Poolside | 37 | 38 | 20 |
| Laguna S 2.1 | Poolside | 35 | 36 | 19 |
| Nemotron 3.5 Content Safety | NVIDIA | 44 | 0 | 0 |
| Gemma 4 31B | 12 | 11 | 0 | |
| Gemma 4 26B-A4B | 0 | 0 | 0 |
“Readable” is stricter than “answered”: a reply with no parseable option is a reply, and it is not an answer to the question that was asked. 212 asks were refused by the free tier's rate limiter rather than by a mind, and are recorded as that and never scored.
Does “I am not sure” mean anything?
The open prompt ends by inviting a hedge: if you are not sure, give your best answer anyway and say that you are unsure. So the RATE of hedging here is not a measurement of spontaneous doubt, and is not reported as one. What is worth knowing is whether the hedge is informative: when one of these minds says it is unsure, is it more often wrong? The marker list is narrow, published in analyse.mjs, and deliberately excludes “I think” and “I believe”, which are register rather than calibration.
| open answers | reader A | reader B |
|---|---|---|
| that said outright they were unsure | 15 | 15 |
| of those, correct | 60.0% | 60.0% |
| that did not | 374 | 374 |
| of those, correct | 74.6% | 74.6% |
| Fisher's exact, two-sided | 0.2313 | 0.2313 |
Its denominator is not the headline's. This table is over every open answer, including the ones marked as no answer at all; the headline is over pairs, which exist only where the same mind also gave a readable option choice. That is why the plain-answer accuracy here sits below the 87.0% in the table above, and the two are not in conflict: one counts a guardrail model's safety classification as a wrong answer and the other has nothing to pair it with.
This is descriptive and it is the smallest cell on the page, so read it with its counts in view rather than its percentages. It is here because the data was already in hand and the question is worth a number: a hedge that does not predict an error is a hedge that tells a reader nothing.
The audit this experiment owes itself
Every checked answer here is one this project published and showed its working for. If one of our pages is wrong, the item is wrong, and the minds disagreeing with us is the only signal available from inside. So the pre-registration fixed a rule before the run: any question where fewer than a third of the answering minds give our answer in both formats is flagged, whatever it does to the headline.
1 question(s) flagged
- When air splits at the front of a wing, what happens to the two halves at the trailing edge? · our answer: They never rejoin; the top air arrives first. Four options: 25% chose it; asked openly: 20%.
A flag is not a verdict against us. It marks a question where the minds and this corpus disagree hard enough that a reader should look at our page and decide for themselves, and the link to each is in the table of every question above.
What this cannot tell you
- The system prompt is confounded with the format, and the control that would have measured it WAS NOT RUN. A multiple-choice prompt must ask for a number and an open prompt cannot, so the conditions differ in more than the presence of options. The registered fourth condition, which removes the open prompt's three-sentence bound, collected zero cells: the free tier's daily ceiling was reached, and the flag it needed had never been deployed. Every gap here is a gap between two whole conditions and nothing here separates the format from the prompt. This is the study's largest weakness and it is not a small one.
- Chance is 25% under four options and undefined under none. Both the raw and the chance-corrected gap are printed for that reason, and where they disagree the disagreement is the finding.
- Four options also tell the mind what kind of answer is wanted. Part of any gap is elimination and part is that the options disambiguate the question. This experiment cannot separate the two, and that is a real limit on reading the gap as “models know less than the benchmark says”.
- Answers within a question are not independent. Every inference here takes the item as the unit; the pooled figures are descriptions.
- One catalogue, one night, one key. These are the minds OpenRouter served free to this project on the day of the run, at temperature 0, which is not how anybody uses them either. Nothing here generalises to paid or frontier models.
- Both readers are instances of the same model family. Every response is published verbatim so that a human can overrule either of them.
The check
Everything above is regenerated from the committed transcripts by research/nobody-gets-four-options/analyse.mjs, and verify-nobody-gets-four-options.mjs re-derives the accuracy rates, the exact sign test, its p-value and McNemar's discordant counts from the raw records in code that does not import the analysis, re-implements the mechanical screen from the committed patterns and requires it to reproduce every verdict, checks every quoted model answer against a transcript, checks every item's evidence quote against the published page it cites, and asks git whether the pre-registration predates the first question. With --browser it also drives this page in a real browser: that the marking panel renders a response and reveals both readers' marks when you pick one, that the both-ways panel renders four options and advances, and that no inline script on the page makes a network call or writes to storage, which is the check behind the sentence promising that nothing you do here is recorded. The transcripts, the marks, the blind sheet and the key are all in the research directory.
What is left open
The sharpest arm this study did not build, specified here so it does not have to be re-derived: the leading question. Both formats above are neutral, and neither is how a person who already believes the myth would ask. Asking “why is it hotter in summer when the Earth is closer to the Sun?” puts the myth in as a presupposition and measures something neither arm here can see: whether a mind corrects a false premise or answers inside it. The rail it needs (POST /api/ask) is built and live; what it needs is a presuppositional phrasing per item, its own two-mark rubric (does the response correct the premise or accept it), and the same adversarial review these items got.