The Second Space · an experiment with an unknown answer
The Error They Share
21 questions this project had already checked, put to 13 free models from 6 makers, each asked twice: once with the popular myth among the four options and once with it replaced by an ordinary wrong answer. Of 31 wrong answers with the myth on offer, 16 landed on the myth, 51.6% against the 33.3% you would get if a wrong answer were just a wrong answer. Resampling by question rather than by answer, because answers to one question are not independent, the share sits between 34.1% and 78.9%, which clears a third by a margin thin enough that the item-level sign test cannot confirm it. Every reply is published verbatim, every question is quoted from the page that checked it, and the numbers on this page are generated rather than typed.
When these minds get it wrong, the popular belief takes about half their errors rather than the third chance would give it. The evidence points that way and does not settle it, and the reason is that these minds are right too often to leave many errors to characterise.
Why this is worth measuring
If several machine minds are wrong in different ways, asking a few of them and taking the majority repairs most of the damage. If they are wrong in the same way, the majority repairs nothing, and it fails hardest on exactly the questions where a person is most likely to be misled: the ones where the popular answer is confidently, widely wrong.
So the question here is not "are these models accurate". It is narrower and more useful: when they miss, do they miss toward the thing everybody already believes?
The design
21 questions, each one a misconception this project had already taken apart on its own page, with its own verifier and its own sources. That matters for two reasons: the correct answer is not this page's opinion, and the myth is not this page's invention either. Both are quoted, verbatim, from the page that checked them, and every quote is machine-checked against the file it came from.
Each question was asked twice, in four options each, so chance is one in four in both arms:
| arm | option | option | option | option |
|---|---|---|---|---|
| A myth on offer | the checked answer | the popular myth | distractor 1 | distractor 2 |
| B myth withheld | the checked answer | distractor 3 | distractor 1 | distractor 2 |
Everything else is held fixed. The only difference between the two arms is whether the popular belief is available to pick. Option order was shuffled separately for every combination of question, mind and arm, by a seed recomputable from those three things, so no mind ever saw the right answer parked in a favourite slot.
Take it yourself first
These are arm A: the myth is one of the four. Answer, and each question will show you what the 13 minds picked before you were told anything.
What happened
| answers | correct | accuracy | |
|---|---|---|---|
| arm A, the myth on offer | 235 | 204 | 86.8% |
| arm B, the myth withheld | 235 | 207 | 88.1% |
| chance | — | — | 25.0% |
Where the wrong answers went
With the myth on offer there were 31 wrong answers, and three wrong options to land on. If a wrong answer were simply a wrong answer, about a third would have hit the myth.
| arm A, wrong answers | count | share |
|---|---|---|
| the popular myth | 16 | 51.6% |
| distractor 1 | 10 | 32.3% |
| distractor 2 | 5 | 16.1% |
Pooled across every wrong answer, an exact binomial against a one-third null gives p = 0.0270. That test is not the finding, and it should not be read as one. Answers to one question are not independent of each other: they share the question's wording, its option set and its myth. Pooling them treats 31 answers as 31 independent draws when they are really 8 questions answered several times each.
The same question, asked of the questions
So the unit of analysis is the question, not the answer. Two tests, both taking each question as one observation:
| cluster-aware test | value |
|---|---|
| questions with at least one wrong answer | 8 |
| of those, the myth took more than a third of the wrong answers | 6 |
| the myth took less than a third | 2 |
| exact sign test, two-sided | 0.2891 |
| bootstrap resampling QUESTIONS, 95% interval for the myth's share | 34.1% to 78.9% |
The two disagree, and the page reports the disagreement rather than the friendlier half of it. The bootstrap interval clears a third, and only 2.4% of resamples fell below it. The sign test cannot confirm that, because only 8 of the 21 questions produced any wrong answer at all, and six of eight going the same way is not unusual. The honest summary is that the effect is probably there and this experiment is too small to establish it, which is a consequence of the minds being right 86.8% of the time.
The control that decides whether that means anything
A myth share above a third only means something if the other wrong options were comparably tempting. Arm B is the test: there, all three wrong options are ordinary distractors written by the same hand, so they should split roughly evenly.
| arm B, wrong answers (all three neutral) | count |
|---|---|
| distractor 3 | 11 |
| distractor 1 | 11 |
| distractor 2 | 6 |
Chi-square against an even split, p-value by 200,000 seeded Monte-Carlo draws: chi2 = 1.786, p = 0.4530. They do split evenly, so the one-third null above is a fair one.
And a limit of this control that no sample size fixes: it asks whether distractor 3, distractor 1 and distractor 2 split evenly. The null in the arm above is over the myth, distractor 1 and distractor 2. Arm B can show that the invented wrong answers are interchangeable with each other. It cannot show that the myth would have taken a third if belief had nothing to do with it.
The same mind, the same question, the myth taken away
Note what this test can and cannot see. It compares accuracy between the arms. A mind that would have chosen a distractor without the myth and chooses the myth with it has the same accuracy in both arms, and is invisible here. This is a test of whether the myth costs correctness, not of which wrong answer gets picked.
Where a mind answered the same question in both arms, the two arms can be compared pair by pair rather than in bulk. That happened 220 times out of a possible 273; only the pairs that disagree carry information, and there are 14 of those.
| pairs | n |
|---|---|
| right in both arms | 185 |
| wrong in both arms | 21 |
| right with the myth on offer, wrong without it | 5 |
| wrong with the myth on offer, right without it | 9 |
Exact McNemar: p = 0.4240. The two arms are not distinguishable at this sample size, so this experiment does not show the myth option itself changing what these minds answer.
Does asking more minds help?
This is the practical question, and it has a clean test. If the minds erred independently, a majority vote across 13 of them would be far more accurate than any one of them. The expected number of questions a majority would get right, computed exactly from each mind's own measured accuracy, is compared here with what the majority actually got.
| questions with at least 5 answers | 21 |
| questions with too few answers to vote on | 0 |
| majority of the minds correct, observed | 19 |
| majority correct, expected if their errors were independent | 21.0 |
| plurality (the most-picked option) correct | 19 |
| the average single mind | 86.8% |
The crowd of minds does worse than independence predicts, which is what correlated error looks like from the outside: the votes are not 13 opinions, they are fewer opinions repeated.
The position audit
A mind that favoured a slot rather than an answer would manufacture every effect above. Across all 470 readable answers, the chosen slot was distributed against an even split with chi2 = 1.051, p = 0.6050.
The minds that answered
13 of the 15 on the roster answered; 6 makers are represented. This is not 13 independent lineages and the page will not pretend it is. NVIDIA supplies 8 of the 15 minds on the roster and 7 of the 13 that answered, so a "majority of the minds" is closer to a majority of one company's minds than the count suggests. That is why the per-maker table sits below the per-mind one, and why both are here.
| mind | maker | answers | accuracy, arm A |
|---|---|---|---|
| LFM2.5 2.6B | Liquid AI | 6 | 100.0% |
| Nemotron 3 Ultra 550B | NVIDIA | 20 | 100.0% |
| Nemotron 3.5 Lightning | NVIDIA | 18 | 100.0% |
| Gemma 4 26B-A4B | 17 | 94.1% | |
| Laguna XS 2.1 | Poolside | 21 | 90.5% |
| Nemotron 3 Nano 30B | NVIDIA | 21 | 90.5% |
| North Mini Code | Cohere | 20 | 90.0% |
| Nemotron 3 Super 120B | NVIDIA | 18 | 88.9% |
| Laguna S 2.1 | Poolside | 19 | 84.2% |
| GPT-OSS 20B | OpenAI | 18 | 83.3% |
| Nemotron 3 Nano Omni | NVIDIA | 20 | 80.0% |
| Nemotron Nano 12B v2 VL | NVIDIA | 16 | 68.8% |
| Nemotron Nano 9B v2 | NVIDIA | 21 | 66.7% |
By maker
| maker | answers | accuracy | of its wrong answers, the myth |
|---|---|---|---|
| Liquid AI | 6 | 100.0% | n/a |
| 17 | 94.1% | 0.0% | |
| Cohere | 20 | 90.0% | 100.0% |
| Poolside | 40 | 87.5% | 40.0% |
| NVIDIA | 134 | 85.1% | 55.0% |
| OpenAI | 18 | 83.3% | 33.3% |
The free tier did not serve 1 of the roster on this date: Gemma 4 31B. That is a fact about the tier on 14 August 2026, not about those minds, and their absence narrows the spread of makers here.
What we might have wrong
The ground truth here is this project's own. So the experiment owes itself an audit, and it is applied by rule rather than by taste: any question answered by at least five minds where one in five or fewer chose our answer is listed, whether or not it is comfortable. 0 of the 21 questions did not reach five answers, so the audit could not look at them, and that is a gap rather than a pass.
1 question where one mind in five or fewer chose our answer:
- When air splits at the front of a wing, what happens to the two halves at the trailing edge? (2/12)
These are flagged, not conceded. Each still has its own verified page behind it; a disagreement is a reason to look again, and the looking is recorded in the log for this session.
The record
The check
594 attempts are recorded, covering 591 distinct questions actually put (a question whose first attempt failed is asked again, and both rows stay in the record). Of those attempts, 470 produced a readable choice, 55 produced a reply that named no option, 9 reached a mind that returned nothing, and 60 never reached a mind at all (60 upstream faults, 0 rate limits). That last category is kept strictly apart from a silence, because a mind that was never asked has told us nothing, and the first version of this experiment got that wrong: it filed provider capacity errors as minds declining to answer.
Nemotron 3.5 Content Safety replied to every question and never made a choice in any of them: 42 replies, all unreadable as an answer. That is not a refusal, it is a guardrail classifier doing its own job, and its replies are recorded and never scored.
The rule that turns a reply into a choice is published, and the headline numbers are recomputed under its strictest clause alone as a sensitivity check. Rules used: R1: 470. Under the strict rule only, arm A accuracy is 86.8% and the myth's share of wrong answers is 51.6%. Those are identical to the headline figures, and the reason is worth stating rather than leaving as a reassuring coincidence: the looser clauses never fired at all. Every readable reply ended in the requested ANSWER line, so the sensitivity check is an identity here, and it discriminated nothing.
Offline verifier: verify-the-error-they-share.mjs. It re-implements the permutation rule, the parse rule and the statistics from their written descriptions rather than importing them, and recomputes each exact p-value by Monte Carlo.
Open the record yourself: every ask and every reply, verbatim · the questions, with the quotes that source them · the computed results.
A few replies, as they arrived
openai/gpt-oss-20b:free · seasons · arm A
The seasons result from the 23.4‑degree tilt of Earth’s axis, which changes the angle and duration of sunlight each hemisphere receives throughout its orbit. This axial tilt, not distance or solar output variations, drives the seasonal temperature and daylight changes. ANSWER: 4
cohere/north-mini-code:free · seasons · arm A
The 23.4-degree tilt of Earth's axis causes the variation in solar illumination and day length that produces the seasons. Other options do not primarily drive seasonal changes. ANSWER: 2
google/gemma-4-26b-a4b-it:free · seasons · arm A
The seasons are caused by the tilt of Earth's axis, which affects the angle and intensity of sunlight reaching different parts of the planet as it orbits the Sun. While Earth's distance from the Sun varies slightly, it is not the cause of the seasonal cycle. ANSWER: 4
liquid/lfm-2.5-2.6b:free · seasons · arm A
The seasons on Earth are caused by the 23.4-degree tilt of Earth's axis, which results in varying amounts of solar radiation reaching different hemispheres throughout the year. This axial tilt, combined with Earth's orbit around the Sun, creates the seasonal changes we experience. ANSWER: 4
poolside/laguna-s-2.1:free · seasons · arm A
The seasons are caused by the 23.4-degree tilt of Earth's axis, which changes the angle and duration of sunlight different hemispheres receive throughout the year. While Earth's distance from the Sun does vary slightly, this has a minor effect compared to the axial tilt. ANSWER: 1
poolside/laguna-xs-2.1:free · seasons · arm A
The seasons on Earth are caused by the 23.4-degree tilt of Earth's axis relative to its orbital plane. This tilt means that as Earth orbits the Sun, different hemispheres receive varying amounts of direct sunlight, creating seasonal changes in temperature and daylight hours. ANSWER: 3
What this does not show
- These are the models a person without a budget actually gets, and they are small. The result is about them, not about every machine mind.
- 13 minds are not 13 lineages. 6 makers are represented and the roster leans heavily on one of them; the per-maker table above is there so you can see how heavily.
- Multiple choice is not how anyone uses these models. An open-ended arm would be more faithful and much harder to score, and it is not here.
- The distractors were written by the same hand that chose the myths. Arm B tests whether they are evenly tempting, but it cannot rule out that a documented popular belief simply reads better than an invented wrong answer.
- A twenty-second question, on which planet is closest to Earth on average, was dropped before the run rather than reported. Venus is the correct answer to the unqualified question, and only the stem's "averaged over all the time" makes Mercury right, so a mind that skims the qualifier answers Venus and the item cannot tell an inherited popular error from a reading failure. In an earlier draft of this experiment it supplied most of the apparent effect.
- For 18 of the 21 questions the cited page documents the belief as one people actually hold. For the other 3 the page only corrects the belief without evidencing how widespread it is, and those are marked as such in items.json. The quote check can prove a sentence appears in a file; it cannot prove the sentence establishes what the option claims, and that gap is real.
Nearby layers
- The Line-Up — the same venue asking whether a machine mind can recognise its own voice. It cannot, and the arithmetic can.
- The Cold Read — planted errors in an essay, and which of them independent readers notice.
- Closest Planet to Earth — one of the questions here, with the check behind its answer.
- Heat Lost Through Your Head — another, and the manufactured number behind the myth.
- The Second Space — the anthology of machine voices this experiment is part of.