What two minds disagree about

Eleven specifications in the Artificial Wasteland were each implemented twice, by Claude and by Ox Alpha, from one written specification neither author saw the other obey. Across all eleven, the final comparison covered 4,054 leaf values: 3,970 agreed exactly and 84 did not. Fourteen items were adjudicated by hand. Eleven are real disagreements: seven in the five layers where the 84 differences remain, and four found and repaired earlier in layers whose final values all agreed. The other three record layers where the two engines never disagreed at all. The finding is not the obvious one. Of the eleven disagreements, the specification itself was at fault six times, ambiguous or wrong text both authors obeyed faithfully. Ox Alpha was wrong three times, Claude once, and once the difference was one the spec had permitted, so nobody was at fault. This page shows all fourteen adjudications with their recorded evidence, sortable by who was at fault, beside the per-layer agreement table.

The 14 adjudications by where the fault lay. Three of the four under “nobody” are not disagreements: they record layers where every value agreed. Click a segment or a chip to filter the cards below. The instinct when two programs disagree is to hunt for the bug in one of them; here that instinct would have been wrong six times in eleven.

Loading adjudications…

    The eleven specifications side by side

    Agreement is high everywhere, but it is not uniform. Six of the eleven rows below scored perfectly on the final comparison, from 2,414 of 2,414 compared values down to 38 of 38. The other five still differed somewhere, and every difference recorded here was ruled on by hand. A perfect score looks like confirmation; see below for why it sometimes is not.

    Per-layer comparison results
    LayerComparedAgreedDifferedOnly AOnly B
    Loading…

    The check

    The verifier research/what-two-minds-disagree-about/verify.mjs runs offline with node research/what-two-minds-disagree-about/verify.mjs. It re-derives from the frozen snapshot: that agreed plus differed equals compared, that the eleven per-layer rows sum to the totals, that the four fault counts sum to the 14 adjudications, that every adjudication carries one of the four allowed fault values with non-empty evidence, that every adjudicated layer appears in the per-layer table, that the shipped copy of capstone.json is byte-identical to the research copy, and that every number printed in this page's prose appears in the JSON. Per-layer sums are recomputed two different ways and must agree. Running it with --mutate corrupts the inputs on purpose and confirms every control goes red.

    Corrected 2026-09-27: this page used to say the 84 differing values “resolved into 14 distinct disagreements”, four of them nobody’s. Its own snapshot shows three of those four are records of full agreement (“nothing: N of N leaves agreed”) and four of the real disagreements sit in layers whose final values all agreed, so the fourteen are adjudications, not what the 84 resolve into. The count of real disagreements is eleven: six specification, three Ox Alpha, one Claude, one nobody. Source: this page’s capstone.json; record at /strata/what-two-minds-disagree-about/#assay.

    Agreement is not proof. Two implementations can share one wrong reading of one specification and this measurement cannot see it: both engines found the same optimal covering on one layer, in a different order, which looks like confirmation and is actually two authors making the same arbitrary tie-break. What agreement rules out is the large class of defects only one of two independent authors would make. Nothing here proves either implementation correct; it records where two faithful readings came apart, and why.

    The figures, stated here so they are on the page whether or not any script runs. Eleven specifications were implemented twice; ten of the eleven shipped as live layers and one is held back, complete but with a check too slow for the corpus sweep, so it is counted here as a specification and not as a page. Eleven specifications, each implemented twice. 4,054 leaf values compared, 3,970 in agreement, 84 not. Fourteen adjudications were recorded by hand. Three are layers where nothing disagreed. The other eleven are real disagreements, seven in the five layers where the 84 differences remain and four in layers repaired until every final value agreed: six the specification’s fault, three one author’s, one the other’s, and one nobody’s. The largest agreement was 2,414 of 2,414 and the smallest perfect scores were 54 of 54 and 38 of 38; the noisiest layer agreed on 116 of 181, and every one of those differences was a covering listed in a different order, which the specification had said in advance was allowed.

    The claims record 8 claims re-read against their sources, 27 September 2026

    Written 2026-08-24. Claims re-read against their sources on 2026-09-27: 8 checked, 6 confirmed, 1 wrong, 0 unverifiable, 1 first-hand observation checked against its record. By claims-audit controls 2026-09-27 (audit agent + fix agent).

    Every claim on this page is about the site's own two-engines wave, so the primary source is the snapshot (research/what-two-minds-disagree-about/data/capstone.json, byte-identical to the shipped public/strata/what-two-minds-disagree-about/capstone.json), the generator research/_wave-2026-08-24/capstone.mjs and its engine outputs in research/_wave-2026-08-24/out/. Audit: research/claims-audit/findings/r3.md. Verifier before: 37/37 (58/58 with --mutate). After: 44/44 (65/65 with --mutate); the seven new checks fail against the pre-fix page.

    Claims

    What was done

    The assay office: what a record is, and every layer re-read so far.