Artificial Wasteland / mind

A living experiment in six tests

The test that tests a set

In Wason's 1960 published aggregate, 6 of 29 participants made a correct first announcement. Try six tests, inspect the exact anonymous payload, then watch an anytime-valid living record grow.

A positive test is not the same thing as positive feedback, and neither is proof of a wish to be right. The geometry between a guessed set and the target set decides what a test can expose.

First, do the thing

One example. One hidden rule.

The ordered triple 2, 4, 6 fits a rule. Enter up to six triples. Before each result, say whether your triple was meant to match your current guess, not match it, or whether you had no current guess. Then make one fixed-menu announcement.

Have you seen a hidden-rule task beginning 2, 4, 6 before?

Three unlike populations

Comparable columns, not interchangeable samples

The dead arm is a copyrighted paper's printed aggregate. The living arm is a self-selected web counter under a shorter fixed-menu task. The machine arm contains independent first actions, not full interactive episodes. Their protocols differ, so no honest human-versus-machine rank follows.

Arm A / the dead / aggregate only

Wason, 1960

6 / 29

Correct at the first announcement, derived from the transcribed Table I announcement matrix and agreeing with the four cells printed in the paper's abstract. This is an historical proportion in 29 psychology undergraduates, not a human constant.

Table I, Frequency of announcements, p. 132

AnnouncementCorrectNoneIncorrect

Each incorrect count carries forward to the next announcement; each none count made no further announcement. Only 10 of the 13 subjects who reached one incorrect conclusion were correct at their next announcement; three announced nothing more (p. 138).

Abstract distribution, p. 129

Printed categoryPrinted cellDerived from Table IRecomputed share
Cell sum / printed N

... / ...

Correct / incorrect / none

... computed from Table I
... printed

Derived eventual outcome
... Derived from the Table I columns, not printed as a single figure; the paper's p. 138 text is the cross-check.

Checking the transcription before opening the living arms...

Arm C / machines / first action only

Famous surface, affine surface

Looking for the supervisor's raw completions...

The affine seed 17, 23, 29 removes the famous surface string while preserving the same equal-step example. It does not prove absence from training. The contrast is a contamination diagnostic, not a clean measure of memorization.

Arm A, reproduced rather than decorated

What the paper printed, and what it did not

The article prints aggregate cells and six example protocols, not all 29 participant records. This page does not expand four counts into invented rows. The group values below are transcribed summaries, displayed as printed rather than re-estimated.

Printed group summaryCorrect first announcement, n=6Incorrect first announcement, n=22Source
Mean incompatible:compatible instance ratio before announcement......p. 133
Mean actual-negative-instance proportion before announcement......pp. 133-134
Printed one-tailed p for each comparison...pp. 133-134

These printed p-values belong to the 1960 analysis. They are not recalculated from unavailable participant records and are not applied to the living arm.

The check

What can break this claim

The analysis seal

The analysis module is fixed by a literal SHA-256, then checked against the served bytes on every load. Recomputing a new hash and displaying it would seal nothing.

ceae58b4fb348edd54f7ffebd5550cfa5e087d2dd4456cad897af324c5620cf6

Checking served bytes...

The aggregate seal

The transcription's sample, the Table I announcement matrix and its row-by-row balance, the abstract's four cells and their derivation from the matrix, the partition, the summaries, the page references, and the canonical checksum must all agree before Arms B and C appear.

checking...

Free analytic choices

The rule menu and six storage classes were fixed before arrivals. A positive test means the reader selected match before feedback. A negative result means the hidden rule rejected a triple. These are different axes. The living display uses 95% betting confidence sequences on a disclosed 401-point grid after 10 in-protocol contributions. No row is screened on its outcome. A stored row that violates the page protocol itself, by claiming more than six tests, classified counts above its own total, or a correctness bit inconsistent with its announced class, is excluded from summaries with its count and reason disclosed beside the living display; this page cannot produce such a row.

The frozen wire forced a narrower study

The store has six aggregate integer fields. It cannot retain assignment, ordered triples, timings, nonces, per-action feedback, or exclusion codes. So this page does not claim the scout's randomized single-versus-dual causal contrast, risk-difference sequence, screening analysis, terminal-history analysis, or an accumulated four-region path map.

What this cannot conclude

Contributors are not a representative population. Familiarity is self-reported. A fixed menu primes and narrows hypotheses. Six tests and one announcement differ from the 1960 administration. The 1 of 29 paper category reached no conclusion is not the same as failing under a web cap.

Interpretive limits

A high positive-test share does not establish irrationality, confirmation seeking, or a desire to be right. Which tests can falsify depends on set geometry. Dual-label facilitation has been observed, but explanations involving complementarity, contrast, information, and descending examples remain contested.

Model limits

The model arm is not novel and not a full interactive replication. The parser counts every completion as valid, refusal, or unparseable. A transformed seed is not contamination-proof. Volunteer and model samples are not placed on one ability scale.

Uncertainties still open

No archive of all 29 record sheets was located, but absence cannot be proved. No formal Crossmark correction status was verified. No affine prompt can be shown absent from a model's training data. The live rate limit deters repetition without proving unique people.

Sources and provenance

Where the checks lead

  1. P. C. Wason, 1960, On the failure to eliminate hypotheses in a conceptual task, pp. 129-140. Numerical aggregate transcription uses pp. 131-134. The article is copyrighted and is not redistributed here.
  2. J. Klayman and Y. Ha, 1987, Confirmation, disconfirmation, and information in hypothesis testing. This supplies the set-relation account of when positive tests can falsify.
  3. R. D. Tweney and colleagues, 1980, Strategies of rule discovery in an inference task. This introduced the dual-label DAX/MED framing discussed here. Its full numerical table was not available to the scout, so unverified values are not printed.
  4. M. Gale and L. J. Ball, 2006, Dual-goal facilitation in Wason's 2-4-6 task. This is evidence for facilitation and for uncertainty about its mechanism.
  5. S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon, 2021, Time-uniform, nonparametric, nonasymptotic confidence sequences. The shared kit documents its betting implementation and guarantees.
  6. E. Banatt and colleagues, 2024, WILT; A. R. Jhaveri and colleagues, 2026, Failing to Falsify; L. Bertolazzi and colleagues, 2026, FALSIFYBENCH. These preprints make any first-LLM-study claim untenable.