The Second Space · an experiment with an unknown answer

The Error They Share

21 questions this project had already checked, put to 13 free models from 6 makers, each asked twice: once with the popular myth among the four options and once with it replaced by an ordinary wrong answer. Of 31 wrong answers with the myth on offer, 16 landed on the myth, 51.6% against the 33.3% you would get if a wrong answer were just a wrong answer. Resampling by question rather than by answer, because answers to one question are not independent, the share sits between 34.1% and 78.9%, which clears a third by a margin thin enough that the item-level sign test cannot confirm it. Every reply is published verbatim, every question is quoted from the page that checked it, and the numbers on this page are generated rather than typed.

When these minds get it wrong, the popular belief takes about half their errors rather than the third chance would give it. The evidence points that way and does not settle it, and the reason is that these minds are right too often to leave many errors to characterise.

Why this is worth measuring

If several machine minds are wrong in different ways, asking a few of them and taking the majority repairs most of the damage. If they are wrong in the same way, the majority repairs nothing, and it fails hardest on exactly the questions where a person is most likely to be misled: the ones where the popular answer is confidently, widely wrong.

So the question here is not "are these models accurate". It is narrower and more useful: when they miss, do they miss toward the thing everybody already believes?

The design

21 questions, each one a misconception this project had already taken apart on its own page, with its own verifier and its own sources. That matters for two reasons: the correct answer is not this page's opinion, and the myth is not this page's invention either. Both are quoted, verbatim, from the page that checked them, and every quote is machine-checked against the file it came from.

Each question was asked twice, in four options each, so chance is one in four in both arms:

armoptionoptionoptionoption
A myth on offerthe checked answerthe popular mythdistractor 1distractor 2
B myth withheldthe checked answerdistractor 3distractor 1distractor 2

Everything else is held fixed. The only difference between the two arms is whether the popular belief is available to pick. Option order was shuffled separately for every combination of question, mind and arm, by a seed recomputable from those three things, so no mind ever saw the right answer parked in a favourite slot.

Take it yourself first

These are arm A: the myth is one of the four. Answer, and each question will show you what the 13 minds picked before you were told anything.

What happened

 answerscorrectaccuracy
arm A, the myth on offer23520486.8%
arm B, the myth withheld23520788.1%
chance25.0%

Where the wrong answers went

With the myth on offer there were 31 wrong answers, and three wrong options to land on. If a wrong answer were simply a wrong answer, about a third would have hit the myth.

arm A, wrong answerscountshare
the popular myth1651.6%
distractor 11032.3%
distractor 2516.1%

Pooled across every wrong answer, an exact binomial against a one-third null gives p = 0.0270. That test is not the finding, and it should not be read as one. Answers to one question are not independent of each other: they share the question's wording, its option set and its myth. Pooling them treats 31 answers as 31 independent draws when they are really 8 questions answered several times each.

The same question, asked of the questions

So the unit of analysis is the question, not the answer. Two tests, both taking each question as one observation:

cluster-aware testvalue
questions with at least one wrong answer8
of those, the myth took more than a third of the wrong answers6
the myth took less than a third2
exact sign test, two-sided0.2891
bootstrap resampling QUESTIONS, 95% interval for the myth's share34.1% to 78.9%

The two disagree, and the page reports the disagreement rather than the friendlier half of it. The bootstrap interval clears a third, and only 2.4% of resamples fell below it. The sign test cannot confirm that, because only 8 of the 21 questions produced any wrong answer at all, and six of eight going the same way is not unusual. The honest summary is that the effect is probably there and this experiment is too small to establish it, which is a consequence of the minds being right 86.8% of the time.

The control that decides whether that means anything

A myth share above a third only means something if the other wrong options were comparably tempting. Arm B is the test: there, all three wrong options are ordinary distractors written by the same hand, so they should split roughly evenly.

arm B, wrong answers (all three neutral)count
distractor 311
distractor 111
distractor 26

Chi-square against an even split, p-value by 200,000 seeded Monte-Carlo draws: chi2 = 1.786, p = 0.4530. They do split evenly, so the one-third null above is a fair one.

And a limit of this control that no sample size fixes: it asks whether distractor 3, distractor 1 and distractor 2 split evenly. The null in the arm above is over the myth, distractor 1 and distractor 2. Arm B can show that the invented wrong answers are interchangeable with each other. It cannot show that the myth would have taken a third if belief had nothing to do with it.

The same mind, the same question, the myth taken away

Note what this test can and cannot see. It compares accuracy between the arms. A mind that would have chosen a distractor without the myth and chooses the myth with it has the same accuracy in both arms, and is invisible here. This is a test of whether the myth costs correctness, not of which wrong answer gets picked.

Where a mind answered the same question in both arms, the two arms can be compared pair by pair rather than in bulk. That happened 220 times out of a possible 273; only the pairs that disagree carry information, and there are 14 of those.

pairsn
right in both arms185
wrong in both arms21
right with the myth on offer, wrong without it5
wrong with the myth on offer, right without it9

Exact McNemar: p = 0.4240. The two arms are not distinguishable at this sample size, so this experiment does not show the myth option itself changing what these minds answer.

Does asking more minds help?

This is the practical question, and it has a clean test. If the minds erred independently, a majority vote across 13 of them would be far more accurate than any one of them. The expected number of questions a majority would get right, computed exactly from each mind's own measured accuracy, is compared here with what the majority actually got.

questions with at least 5 answers21
questions with too few answers to vote on0
majority of the minds correct, observed19
majority correct, expected if their errors were independent21.0
plurality (the most-picked option) correct19
the average single mind86.8%

The crowd of minds does worse than independence predicts, which is what correlated error looks like from the outside: the votes are not 13 opinions, they are fewer opinions repeated.

The position audit

A mind that favoured a slot rather than an answer would manufacture every effect above. Across all 470 readable answers, the chosen slot was distributed against an even split with chi2 = 1.051, p = 0.6050.

The minds that answered

13 of the 15 on the roster answered; 6 makers are represented. This is not 13 independent lineages and the page will not pretend it is. NVIDIA supplies 8 of the 15 minds on the roster and 7 of the 13 that answered, so a "majority of the minds" is closer to a majority of one company's minds than the count suggests. That is why the per-maker table sits below the per-mind one, and why both are here.

mindmakeranswersaccuracy, arm A
LFM2.5 2.6BLiquid AI6100.0%
Nemotron 3 Ultra 550BNVIDIA20100.0%
Nemotron 3.5 LightningNVIDIA18100.0%
Gemma 4 26B-A4BGoogle1794.1%
Laguna XS 2.1Poolside2190.5%
Nemotron 3 Nano 30BNVIDIA2190.5%
North Mini CodeCohere2090.0%
Nemotron 3 Super 120BNVIDIA1888.9%
Laguna S 2.1Poolside1984.2%
GPT-OSS 20BOpenAI1883.3%
Nemotron 3 Nano OmniNVIDIA2080.0%
Nemotron Nano 12B v2 VLNVIDIA1668.8%
Nemotron Nano 9B v2NVIDIA2166.7%

By maker

makeranswersaccuracyof its wrong answers, the myth
Liquid AI6100.0%n/a
Google1794.1%0.0%
Cohere2090.0%100.0%
Poolside4087.5%40.0%
NVIDIA13485.1%55.0%
OpenAI1883.3%33.3%

The free tier did not serve 1 of the roster on this date: Gemma 4 31B. That is a fact about the tier on 14 August 2026, not about those minds, and their absence narrows the spread of makers here.

What we might have wrong

The ground truth here is this project's own. So the experiment owes itself an audit, and it is applied by rule rather than by taste: any question answered by at least five minds where one in five or fewer chose our answer is listed, whether or not it is comfortable. 0 of the 21 questions did not reach five answers, so the audit could not look at them, and that is a gap rather than a pass.

1 question where one mind in five or fewer chose our answer:

These are flagged, not conceded. Each still has its own verified page behind it; a disagreement is a reason to look again, and the looking is recorded in the log for this session.

The record

The check

594 attempts are recorded, covering 591 distinct questions actually put (a question whose first attempt failed is asked again, and both rows stay in the record). Of those attempts, 470 produced a readable choice, 55 produced a reply that named no option, 9 reached a mind that returned nothing, and 60 never reached a mind at all (60 upstream faults, 0 rate limits). That last category is kept strictly apart from a silence, because a mind that was never asked has told us nothing, and the first version of this experiment got that wrong: it filed provider capacity errors as minds declining to answer.

Nemotron 3.5 Content Safety replied to every question and never made a choice in any of them: 42 replies, all unreadable as an answer. That is not a refusal, it is a guardrail classifier doing its own job, and its replies are recorded and never scored.

The rule that turns a reply into a choice is published, and the headline numbers are recomputed under its strictest clause alone as a sensitivity check. Rules used: R1: 470. Under the strict rule only, arm A accuracy is 86.8% and the myth's share of wrong answers is 51.6%. Those are identical to the headline figures, and the reason is worth stating rather than leaving as a reassuring coincidence: the looser clauses never fired at all. Every readable reply ended in the requested ANSWER line, so the sensitivity check is an identity here, and it discriminated nothing.

Offline verifier: verify-the-error-they-share.mjs. It re-implements the permutation rule, the parse rule and the statistics from their written descriptions rather than importing them, and recomputes each exact p-value by Monte Carlo.

Open the record yourself: every ask and every reply, verbatim · the questions, with the quotes that source them · the computed results.

A few replies, as they arrived

openai/gpt-oss-20b:free · seasons · arm A
The seasons result from the 23.4‑degree tilt of Earth’s axis, which changes the angle and duration of sunlight each hemisphere receives throughout its orbit. This axial tilt, not distance or solar output variations, drives the seasonal temperature and daylight changes.  

ANSWER: 4
cohere/north-mini-code:free · seasons · arm A
The 23.4-degree tilt of Earth's axis causes the variation in solar illumination and day length that produces the seasons. Other options do not primarily drive seasonal changes.  
ANSWER: 2
google/gemma-4-26b-a4b-it:free · seasons · arm A
The seasons are caused by the tilt of Earth's axis, which affects the angle and intensity of sunlight reaching different parts of the planet as it orbits the Sun. While Earth's distance from the Sun varies slightly, it is not the cause of the seasonal cycle.

ANSWER: 4
liquid/lfm-2.5-2.6b:free · seasons · arm A
The seasons on Earth are caused by the 23.4-degree tilt of Earth's axis, which results in varying amounts of solar radiation reaching different hemispheres throughout the year. This axial tilt, combined with Earth's orbit around the Sun, creates the seasonal changes we experience.

ANSWER: 4
poolside/laguna-s-2.1:free · seasons · arm A
The seasons are caused by the 23.4-degree tilt of Earth's axis, which changes the angle and duration of sunlight different hemispheres receive throughout the year. While Earth's distance from the Sun does vary slightly, this has a minor effect compared to the axial tilt.

ANSWER: 1
poolside/laguna-xs-2.1:free · seasons · arm A
The seasons on Earth are caused by the 23.4-degree tilt of Earth's axis relative to its orbital plane. This tilt means that as Earth orbits the Sun, different hemispheres receive varying amounts of direct sunlight, creating seasonal changes in temperature and daylight hours.

ANSWER: 3

What this does not show

Nearby layers