A living experiment in projection
How common was the choice you made?
The 1977 table shows people who made opposite choices also gave different estimates of what peers would choose. Make one choice, forecast this page's crowd, then watch differential projection separate from actual calibration as the living reference population grows.
Choose before the crowd appears
Random itemOne low-stakes hypothetical, randomly assigned in this browser session. No result is shown until an accepted send. Your answer has no correct side.
You have already contributed from this browser. The results are open below. Local storage is only an advisory marker, not proof of one person.
Step 1, your choice
Loading the assigned item...
Step 2, your forecast
The thumb is inert until you move it. Whole percentages only. The target is other accepted rows for this exact item, not society.
Step 3, prior exposure
Exact payload, before send
Complete all three steps. Nothing has been sent.
No network write has occurred.
Checking the store without sending...
Three populations, three different claims
The original archive has aggregate cells, the page crowd has accepted rows, and the machine rail has raw completions. They sit beside one another without being pooled.
Arm A, the dead
Four printed rows
Ross, Greene, and House reported 320 Stanford undergraduate responses across four vignettes in 1977. Participant rows were not located. This arm is a transcription of Table 1 aggregates, never a reconstructed dataset.
The four choice-group gaps, each derived from Table 1's printed means, run from 8.9 to 21.5 percentage points. A positive gap says the two choice groups forecast differently. It does not by itself say either group was wrong.
Arm B, the living
The crowd being forecast
Loading accepted rows...
For each item, option-A estimates are compared between people who chose A and B. The same rows also provide a moving ground truth, so each forecast can be compared with all other accepted choices on that item.
Arm C, the machines
Free choices, failures kept
Models receive the same four items under two targets: forecast human visitors, or forecast other completions from the identical model snapshot. No option is forced. If one side has fewer than 10 parsed choices, the difference is not estimable.
The living crowd, item by item
Every interval below is built from the shared betting confidence-sequence kit. A 95% family guarantee means the twelve bounded-mean sequences are valid at every moment simultaneously under their stated stable-mean conditions, even though visitors keep peeking. A familiar 1.96 standard-error interval would only be licensed at a fixed analysis time and could quietly lose its advertised coverage here. Two lines are descriptive and carry no interval: the pooled leave-one-out forecast error and the share of estimates divisible by five. The calibration points are exact leave-one-out values for the received rows, while their paired sequences bound the corresponding population contrast.
The historical anchor, recomputed
Each count and mean below is a published aggregate from Table 1, page 283. Table 1 prints no gap column: the expected gap beside each row is the difference of the two printed means, fixed at transcription time, and the live computed gap from analysis.mjs must reproduce it. Printed F values are listed as reported only because group variances and participant rows are missing.
| story | A n | A group's estimate of A | B n | B group's estimate of A | computed gap | expected gap from printed means | check |
|---|
Reported only: F(1,312) = 49.1, p < .001. Those statistics cannot be recovered from one-decimal means and counts.
The Many Labs 2 preregistration attributed 75.4% and 54.9% to this summary. They are not the values in Table 1. The original table and the final replication article give 65.7% and 48.5%.
The machine rail
A model's human-crowd output is a choice-conditioned forecast about humans, not model false consensus. Same-model runs describe sampled completions, not beliefs or preferences inside a machine. Prior work tested famous scenarios with imposed options after free choices became lopsided; this rail keeps free choices and permits a blank result.
The check
This is the part with sharp edges. The analysis is fixed, the old table is reproduced before the living display can open, and every limitation that changes the claim is kept beside it.
Expected SHA-256: e7244f8696ff53e14c43d065a84ac224858b3c181dd4a7d29981cf5b969db57b
Checking live bytes...
curl -s https://artwaste.land/strata/false-consensus/analysis.mjs | shasum -a 256
The living display is locked if any Table 1 row stops reproducing its expected gap, derived from the printed means, or the printed summary. One-byte seal mutations and wrong-cell anchor mutations are both required to fail offline.
D is mean option-A forecast among A choosers minus the same mean among B choosers. Positive D is differential projection. It is not automatic evidence of error, because a person's own response can be legitimate sample-of-one information.
The accepted crowd supplies a descriptive target the 1977 study did not independently measure. Every calibration number, the two side errors and the pooled forecast error alike, scores forecasts against the other accepted rows only, never against a share that includes the forecaster's own row. These values are exact for these rows, not for society; the paired sequences bound population contrasts under their stated conditions.
The frozen store accepts only item, choice, estimate, and exposure. It has field validation, a per-IP daily limit, a global daily limit, and a row cap. It has no signed assignment token, honeypot, elapsed-time field, or burst flag. Therefore this page has no honest raw-versus-screened split and says accepted row, never screened person.
The browser assignment is random for this session and the local contributed marker is advisory. A cleared browser, another device, coordinated submissions, and bots remain possible. No demographics, free text, account, email, precise time, location, user agent, or fingerprint enters the declared payload.
The twelve streams use alpha 0.05/12 and a 401-point betting grid. Time-uniform validity assumes a stable data-generating mean and predictable bets. It does not make these self-selected visitors representative, remove drift, or repair dependence and gaming.
At zero, the page says zero. Before both choices reach 10 accepted rows on an item, only counts appear. Calibration and D are suppressed. Early ground truth changes every time a row arrives.
No participant-level 1977 data, code, or open data archive was located in the documented search. The one-decimal cells cannot recover distributions, standard deviations, F tests, p values, or participant dots. The displayed F statistic is reported, not reproduced.
This crowd is global, self-selected, online, choice-first, unpaid, and answering one trivial new item. The 1977 participants were Stanford undergraduates answering consequential vignettes, with estimates before choices. Culture, order, item stakes, medium, and familiarity may all change the result.
Only a complete JSON object with exactly two valid keys becomes a value. Refusals and unparseable outputs remain in the denominator, with counts and rates. No regex extraction, repair, reprompt, silent drop, or forced balance is allowed. The result file has no frozen human snapshot, so human-target forecasts cannot honestly be scored against a later moving crowd.
The often repeated meta-analytic r = .31 was not used because the 1985 full results pages were not directly inspected. Its title and publisher record say 115 tests, not the conflicting 155 found in one later reference. No correction or retraction was located for the 1977 or 1985 article as of 2026-08-21, which is a search result, not proof of absence.
Analysis declaration: loading
Sources and exactly what each supports
Ross, Greene, and House (1977): article identity, Table 1 cells, sample description, study order, and printed F statistic. The article is copyrighted; only numerical facts and minimal labels are transcribed.
Klein et al. (2018), Many Labs 2: large preregistered replication and corrected historical summary. Its open participant data are not relabelled as the original and are not shipped on this page.
Dawes (1989): why a self-response can rationally move a prevalence estimate, making the classic relative difference insufficient to prove error.
Waudby-Smith and Ramdas: betting confidence sequences for bounded means. The site kit implements the disclosed predictable plug-in, hedged betting process.
Choi, Hong, and Kim, arXiv:2407.12007v2: prior language-model study, including its famous scenarios and imposed-choice workaround. This page uses new low-stakes items and free samples.
Connections in the Wasteland
The Guess It Showed You First fixes a prediction before a choice. Here the whole analysis is fixed before an accumulating crowd.
Ask a Random Friend shows how the route through a population changes the statistic. Here self-selection and own-choice conditioning define the route.
The Positive Test That's Probably Wrong separates an intuitive estimate from the population base rate. Here the crowd itself becomes the changing base rate.