The Work of Two

Thirteen machine minds were put to twenty-one questions this project had already checked; the free tier being what it is, 11.19 of them answered the average question, and 86.8 per cent of the 235 answers were right. Ask what a panel that size is worth and the answer is not eleven. Measured against a null built by shuffling each mind's own answers, it carries the statistical weight of about two independent minds, and two different routes to that number agree. Most of it sits in a single question about aeroplane wings. Every "eleven" below is that 11.19, the mean number answering, and not a roster.

There is one factor underneath everything on this page, and four fields have four names for it. A survey statistician calls it the design effect. An acoustician meets it when adding voices in a room. A portfolio manager hits it as the floor below which diversification stops working. Condorcet's jury theorem assumes it is exactly one. It is the bracket in this line:

Var(mean of n voices) = (σ² / n) × [ 1 + (n − 1)ρ ]

ρ is the correlation between any two of the voices. At ρ = 0 the bracket is one and the familiar 1/n applies to the variance, so the error of the mean falls like one over the square root of the crowd. At ρ = 1 the bracket is n, the crowd cancels out entirely, and a million voices are one voice said loudly. In between, the useful reading is the effective size, the number of genuinely independent voices that would have produced the same variance:

n_eff = n / [ 1 + (n − 1)ρ ]   →   1 / ρ   as n → ∞

That limit is the part worth staring at. A positive ρ puts a ceiling on a crowd that no amount of recruiting gets past. At ρ = 0.1 an infinite crowd is worth ten. The ceiling is a property of the model, not a measurement, and this page keeps the two apart wherever it prints one.

Instrument one · the factor, and its four names

design effectn/a1 + (n−1)ρ
effective sizen/aindependent voices
ceiling as n→∞n/a1/ρ, a model limit
decibels per doublingn/a3.01 incoherent, 6.02 coherent
Condorcet at p = 0.6n/amajority correct, n_eff jurors
undiversifiable sharen/aof one asset's variance

Every readout is a closed form, computed live from the two sliders by engine/stats.mjs. Nothing here is fitted or simulated.

Why the same bracket keeps turning up

Sound. Take n sources of equal mean-square pressure with pairwise correlation ρ. The total mean-square is n p² times the same bracket, so the level is 10 log10(n[1 + (n−1)ρ]). At ρ = 0 that is 10 log10 n, which is the 3.01 dB per doubling every acoustics course teaches. At ρ = 1 it collapses to 20 log10 n and 6.02 dB. The word to avoid here is the acoustician's own: coherence usually means the magnitude-squared kind, which throws phase away, and two perfectly coherent sources in antiphase cancel rather than adding 6 dB. The bracket needs the signed correlation. The two endpoints are textbook; the line between them is just the algebra.

Money. An equally weighted portfolio of n assets of equal variance and average correlation ρ has variance σ²[1 + (n−1)ρ]/n, which tends to ρσ². The share of a single asset's variance you can never diversify away is exactly ρ, which the instrument above prints.

Surveys. Kish's design effect is the bracket itself: the factor by which a clustered sample's variance exceeds a simple random one, and the number that turns a nominal sample size into an effective one.

Juries. Condorcet's theorem needs jurors who decide independently. It does not degrade gently when they do not: it is a statement about n, and the honest substitution is to put n_eff in its place instead. That substitution is what the last readout above does, and it is an approximation, not a theorem.

None of that is new, and this page is not claiming it. What is new is what happens when you point the bracket at a crowd that actually exists.

The measurement you cannot make

Here is the difficulty that keeps ρ unmeasured almost everywhere the idea of a crowd gets used. You cannot estimate it from one crowd answering one question.

Galton's ox is the founding case. In 1906 he collected 787 legible guesses at the dressed weight of one ox, and this archive has rebuilt that arithmetic. Suppose every guesser was biased upward by the same amount, because the animal looked heavy to everyone. The guesses would spread exactly as they do, and the crowd's median would be off by exactly the shared bias, and nothing in the 787 numbers could tell you which part of the error was shared and which was personal. One question gives you one realisation of the shared component, and one realisation of anything is not a variance.

To separate them you need the same panel answering many questions, so that the shared part can be seen doing the same thing again. That design is ordinary in human judgment research and the decomposition has been done on it (Broomell and Budescu, "Why Are Experts Correlated? Decomposing Correlations Between Judges", Psychometrika 74, 2009, 531; Davis-Stober, Budescu, Dana and Broomell, "When is a crowd wise?", Decision 1, 2014, 79). What is missing is not the method. It is that the quantity almost never gets reported for the crowds people actually rely on, and this project happens to have run the design three times without ever asking it this question.

Three panels, none of them built for this

Between August and September 2026 this project put checked questions to whatever free language models the router would still serve, and published the transcripts. Each of those experiments asked its own question and answered it. None of them asked how much the panel was worth.

One of them came close enough to name the gap and leave it open:

These minds are not independent. They are not, and the roster says so: 15 models from 6 makers, 8 of them NVIDIA, and fewer than that answer on any given night. Results are reported per model and per maker, because "fifteen minds" would be a lie about the sample.

research/myth-consensus/README.md, the objections section, 2026-08-14

That is the seam this page closes. Not are they independent, which the layer already answered, but how far from independent, in the one unit that says what a panel is worth.

Two of the three panels replicate; the third does not, and its failure is worth more than its success would have been. The two forced-choice experiments, run three weeks apart on rosters the free tier had reshuffled, come back at 2.00 and 2.28 effective minds of 11.19, and 2.64 and 2.06 of 7.27 and 6.96, every arm with the null rejecting at the smallest p five thousand shuffles can express. The Line-Up, which asked fourteen minds to name the author of an unattributed passage, comes back at ρ = 0.06 with a 95 per cent interval of −0.20 to 0.30 and p = 0.29. Nothing. Each passage there went to about two judges, so there is almost no power in it, and a panel at chance accuracy is exactly where you would expect to find least.

How this statistic is fooled, and how it fooled this page

The first version of this portal reported that third panel at ρ = 0.393, p = 0.0002, and said in its own frontmatter that the judges were "wrong together as reliably as the high-accuracy panels are". That was false, and the cause is instructive enough to leave on the page rather than in a commit message.

The Line-Up's committed rows carry two different tasks under one field. Most are anthology passages, chance one in five, and the minds score 17.8 per cent on them. The rest are that layer's declared positive control: passages made by asking each mind to state its own model name and its maker, which the judges get right 68.4 per cent of the time. Pooling a give-away task with a chance-level one produces a set of questions that differ enormously in difficulty, and a difficulty spread is precisely the shape this statistic reads as minds failing together. It would have produced a large, significant ρ whatever the minds did.

So the failure mode is general: this measurement is only meaningful over questions that belong to one task. Nothing in the checks here caught it. An adversarial pass over the finished page did, which is the argument for running one.

Instrument two · the panels, and their own null

 

cells scoredn/a 
accuracyn/aof readable answers
dispersion Tn/a 
design effectn/aT over its null mean
ρn/a 
effective sizen/a 

computing…

The permutation null for this panel, with the observed value marked.

What the statistic is, exactly

Every cell is one model's answer to one question, scored 1 or 0 by the parse rule those layers published and share (research/myth-consensus/score.mjs: take the last ANSWER: k, else a bare trailing integer, else an exact match on an option shown; anything else is unreadable and is dropped, never scored as wrong).

The null is the one those layers already use for their own ensemble arithmetic: every model errs independently, at its own measured rate. Under it, the number correct on a question has mean equal to the sum of the answering models' rates and variance equal to the sum of p(1−p). T is the observed squared departure over that variance: one under the null, larger when the minds fail on the same questions.

T is then calibrated against a permutation that shuffles each model's own answers across the questions it answered. That preserves every model's accuracy, every model's answered set and the whole shape of the missing data, and destroys only the alignment between models, which is the thing being tested. The design effect is T divided by the mean of that null, ρ follows from the bracket, and the interval is a bootstrap over questions, which is the unit these experiments were designed around.

One bookkeeping choice moves ρ a long way, so here it is in the open. A model whose rate is exactly 1 contributes nothing to the observed departure and nothing to the variance: it is invisible to T. Three of the thirteen answered everything they were asked correctly. They are nonetheless counted in the 11.19, which is what turns the design effect into a ρ. Count only the models that could have contributed and the mean is 9.10 rather than 11.19, and the same design effect reads ρ = 0.572 and an effective size of 1.62 instead of 0.452 and 2.00. The figures printed here are the first, which is the more flattering of the two for the panel and the more conservative for this page's claim. And note where the mean pairwise correlation in the maker table sits: 0.603, computed over the pairs that have one. Three numbers for one quantity, and the reader is entitled to all three rather than the one that reads best.

What the interval is, and what it is not. The bootstrap resamples questions, which is the right cluster, and on synthetic panels of this shape and size its nominal 95 per cent covers about 85. It is also not a smooth band: 35 per cent of the resamples happen to contain no copy of the wing question and have a median effective size of 3.04, while the rest sit near 1.83. The published interval is a mixture over whether the hardest question survived the draw, which is the same finding as the section below in a different costume.

The check

Re-deriving the panels from the raw transcripts reproduces the members' own published accuracies exactly, which is how you know the same cells are being read: The Error They Share publishes 0.8680851063829788 for the arm with the myth on offer and 0.8808510638297873 for the arm without, and this page's independent pass over the same transcript gets both to the last digit. Nobody Gets Four Options reconciles the same way, cell for cell, against its own per-model readable tallies.

A harder check, because accuracy is a sum and could match by luck: that layer also publishes how many questions a majority of its minds would have got right had they erred independently, computed by an exact Poisson binomial over their individual rates. It publishes 20.96448398614373 out of 21, against 19 observed. Recomputing it from these cells gives 20.964483986143733, agreeing to fourteen decimal places.

An earlier draft called that "a convolution written from scratch here", and an adversarial pass pointed out that it was the member's own recursion with a loop variable renamed. Two runs of the same algorithm on the same input agree to machine precision for reasons that have nothing to do with the data, so that check was testing the bookkeeping and nothing else. It is now also computed a second way that shares only the definition: sum the probability of each of the 213 patterns of who gets it right and bin them by how many did. No recursion, no convolution, 8,192 terms per question. It lands on the same figure, and now the agreement means something. Everything else on this page is computed from those cells in front of you, by engine/stats.mjs, at a fixed seed, so the verifier in research/the-work-of-two/ and your browser agree digit for digit.

The second route, and where it agrees

A design effect is about the variance of a mean, and nobody actually averages a panel of models: they take the majority. So here is the same question asked in the currency people use. Take a random k of the models, score their majority vote, and do it again for every k. Then do the same on permuted copies, which is what this panel would look like if its minds were independent. The number of independent voices whose majority would match what the real panel's majority scored is an effective size read off the vote instead of the variance.

This curve runs on the complete rectangle inside each panel rather than the ragged whole, and it has to. On ragged data a k-subset can only be scored on questions all k of its members answered, and at the largest k only the questions everybody reached survive, which are the easy ones. The ragged curve would climb partly because its questions were getting easier, which is exactly the reassuring artefact this page is testing for.

Instrument three · what the vote is worth

 
observed at full paneln/a 
if they were independentn/asame panel, shuffled
effective size, from the voten/a 
effective size, from the variancen/aon the same rectangle

 

On the headline panel's rectangle the two routes land at 2.16 and 2.57 out of eight, which is as close as two estimators of different things have any business being. Read honestly, that 2.57 carries a convention inside it: the independence curve is jagged because an even panel can tie, and a tie is scored as a half here, so the interpolation happens across the k = 2 to k = 3 step where all the movement is. Take only the odd panel sizes, the ladder Condorcet is usually stated on, and the same observed accuracy reads 2.15 instead. That is a 16 per cent move on an arbitrary choice, and it happens to land the two routes almost on top of each other rather than merely near, which is the direction that should make you more suspicious rather than less.

On two of the four arms the vote route has no resolution at all, because both curves reach 100 per cent and the question stops having an answer; the instrument says so rather than interpolating. On a third it returns 1, clamped at the bottom of the curve. So there is exactly one panel here where both routes produce a number, which is why this is a corroboration and not a replication.

And on one arm the observed majority falls as the panel grows, from 93.6 per cent at one voice to 90.9 at seven. That is Condorcet running backwards on real data, on eleven questions, which is few enough that it should be read as an illustration rather than an estimate. The layer next door simulates that collapse with a conformity model. Here it simply happened. (The formal results on correlated jurors are Ladha, "The Condorcet Jury Theorem, Free Speech, and Correlated Votes", American Journal of Political Science 36, 1992, 617, and Berg, "Condorcet's Jury Theorem, Dependency among Jurors", Social Choice and Welfare 10, 1993, 87. Under bounded exchangeable correlation the theorem does not simply break: majority accuracy still rises with n, but converges below one. This arm is not a counterexample to that, it is eleven questions.)

Where the agreement lives

This is the part that changes what the number means. The correlation is not spread evenly over the questions. It is concentrated, and you can pull it out by hand.

Instrument four · take a question out of the set

Each button is one question, with the share of answering models that got it right. Click to remove it and the panel is measured again from scratch, same seed, same 5,000 permutations. Removing nothing reproduces the headline exactly.

questions keptn/anothing removed
ρn/a
effective sizen/a 
p, against the nulln/a5,000 permutations

 

Remove the questions in order of difficulty and the same thing happens in all four forced-choice arms, and the first one removed is the same question every time. Four arms is two experiments, and they overlap heavily: the later experiment's forty-four questions are a superset of the earlier one's twenty-one, and eight of its nine models are the same models. So this is one finding seen twice under two framings, not four independent sightings.

Effective panel size as the hardest questions are taken out, one at a time.
panelas runless the hardestless twoless three

The question is equal-transit, from The Air That Got There First: when air splits at the front of a wing, what happens to the two halves at the trailing edge. The checked answer is that they never rejoin and the top air arrives first, which is Babinsky's smoke photographs (Physics Education 38, 2003, 497) and NASA Glenn's own list of incorrect lift theories. Two of twelve minds got it. Remove that one question and the eleven go from the work of two to the work of three. Remove it and the metabolism question and they are worth five.

One qualification the option text does not carry, and should. The ordering follows the circulation, and the clean statement is a thin-aerofoil, potential-flow one: for a lifting aerofoil the upper flow arrives first, at exactly zero lift the transit times are equal (which is the myth, true only in the case nobody flies in), and Bai and Wu show the difference can change sign for a very thick section at large angle of attack (Chinese Journal of Aeronautics, 2021, doi 10.1016/j.cja.2021.03.034). Both NASA's sentence and Babinsky's own figure caption keep the word "lifting"; the question here drops it. For a wing in flight the checked answer is right and the myth is wrong, so the finding stands, but a mind that answered against us on the thick-section ground would not be flatly wrong, and a page whose largest single effect rests on this item should say so.

The layer that ran the experiment had already flagged this question by its own rule, which is that any item where the minds overwhelmingly reject our checked answer is a place where we might be the ones who are wrong. It is the only flag in the run. We are fairly sure we are not wrong here, and the sources are on that layer rather than asserted on this one, but the flag is the right instinct and it is worth saying that a portal whose finding rests on one question inherits that question's risk.

Failing together is not agreeing together

On the wing question the twelve answers split five, three, two and two across the four options. They were not wrong in the same way. They were wrong at the same time, which is a different thing and the one the design effect measures. The plurality still landed on a wrong option, so the crowd was no help either way, but a page that said "they all give the same wrong answer" would be describing something the data does not show. The share of errors landing on the popular myth specifically is the question the member layer exists to answer, and it answers it there.

And the five, the largest group, did not go to the myth. They went to "they rejoin a little behind the wing, in the wake", which is a distractor with a real reading behind it: the two streams do merge into a common wake shear layer downstream, they simply do not arrive together and those particular parcels never pair up again. So part of the co-failure that carries this page's headline is sitting on an option that is arguably half right, which is an item-construction explanation for the effect rather than a finding about minds, and it belongs here rather than in a footnote.

What it is not: the makers

The obvious explanation is lineage. Seven of this panel's thirteen models came from one maker; models trained by the same lab on overlapping data should surely fail together more than strangers do. If that were the story, a panel bought from different companies would be a real fix. (The blockquote above says eight of fifteen, which is the source experiment's full roster. Fewer answered on the night, and an earlier draft of this paragraph carried that eight across to a panel it does not describe.)

It is not the story, or at least this data cannot find it. Take each pair of models, take the correlation of their correctness over the questions they both answered, and split the pairs by whether the two models share a maker. Shared question difficulty inflates every pair equally, so the within-minus-across gap is the part difficulty cannot explain. The null shuffles the maker labels among the models, which holds the panel and every pairwise correlation fixed and moves only the grouping.

Not every pair, and the exception matters. A correlation is undefined when one of the two never varied, and three of the thirteen answered everything they were asked correctly, so 33 of the 78 possible pairs have no number and are dropped. The test below therefore runs on the ten models that made at least one mistake, of which five are from the one maker. That is a selected sub-panel, selected on performance, and it is the sub-panel the row counts in the table describe.

panelsame makerdifferent makersdifferencep

Three of the four arms put the difference on the wrong side of zero, and no arm comes near significance. State the power honestly: the headline panel has eleven within-maker pairs against thirty-four across, and one arm has two. This is a null with little power behind it, not a demonstration that lineage does not matter. What it does rule out is the easy version of the story, in which the correlation is obviously and visibly a family resemblance. On this data it is not visible at all.

What is left is the questions. Some questions are hard for everything that answers them, and a panel is only as independent as its hardest question lets it be.

The sentence this was for

On the questions where every mind is right, the panel's agreement costs nothing, because there was no error for a second opinion to catch. On the questions where minds fail, they fail together. A crowd is a real crowd exactly where you do not need one.

An earlier draft of that sentence said the panel was statistically independent on the easy questions, and it was wrong by this page's own arithmetic. Under a null where each mind errs at its own rate, a question everybody gets right is the largest positive departure available: the thirteen unanimous questions here carry a local dispersion of 1.86, well above the independence value of one. They are the most agreeing questions in the set after the two hard ones. What is true is the second half, that the agreement buys nothing, and the first half was a nicer-sounding claim that no statistic on this page supports.

Which makes the effective size a property of the panel and the questions, not of the panel. There is no number you can carry from one use to the next, and you cannot know in advance which of your questions is the aeroplane wing, because if you could you would not have needed to ask.