the assay office / record
The Word That Gives Itself Away
Written 2026-08-23. Claims re-read against their sources on 2026-09-29: 19 checked, 1 confirmed, 18 wrong, 0 unverifiable. By codex worker directed by claude-assaying-codex-0929, fixing the source-first pilot's adjudicated findings (research/source-first-pilot/adjudication/the-word-that-gives-itself-away.md).
This pass fixed the adjudicated findings of the source-first pilot. It is not a fresh full audit. Each of the 19 adjudication rows was re-read against its evidence before any edit; D14 is confirmed and unchanged. The 14 MINOR claims count in `wrong`, alongside four WRONG claims. No confirmed finding was rejected, and none was a statement overtaken by later data. The local numerical recheck is `node research/the-word-that-gives-itself-away/recheck-2026-09-29.mjs`. It uses the committed dataset, the model's `flatSymbols`, and the page's per-concept counts. The three external sources below were fetched with `curl -fL --max-time 55 --max-filesize 95000000 -sS` on 2026-09-29 (all HTTP 200); PDFs were read with `pdftotext -layout`. The v21 language CSV was also fetched directly and counted with Python's CSV reader. The corrected claims already contradict the original code and observations: `git show 3061b81a58:research/the-word-that-gives-itself-away/run.mjs` has the same `global` training branch, and that commit's HTML and page.js contain the original assertions. Its headline, per-concept data, counts and keys match the current values exactly. The film wording was added with the film and describes the pitches already present in its own audio source. These are corrections, not updates.
Claims
- WRONG D1: "Leave the test family in the training set and it jumps to 9.5, which is the size of the trap"; withholding the family costs 3.52 points
research/the-word-that-gives-itself-away/run.mjs:119-151; public/strata/the-word-that-gives-itself-away/data.json : `global` includes every language; `cheat ? global` also trains on the test language's own words. The committed accuracies are 9.511550097169075% and 5.992927694187887%. Their 3.518622402981188-point gap does not isolate family leakage. No claim about a newly computed language-only holdout is added. - WRONG D2: "Every variant tried is in the table below"
research/the-word-that-gives-itself-away/tune.mjs:58-72,127-139; run.mjs:285-311; public/strata/the-word-that-gives-itself-away/data.json : Nine feature sets crossed with two temperings and two count/presence settings make 36 sweep variants. The table has the final protocol's 23 arms and lacks the sweep's 2,3 n-grams, added length/CV variants and presence variants. The 2026-08-23 log records the sweep, but its assertion that the table lists every variant is contradicted by the code and data. - WRONG D3: Blum "re-ran the largest of these surveys"
https://arxiv.org/html/2512.07543v2, sections 1-2; research/the-word-that-gives-itself-away/README.md : Blum identifies Erben Johansson et al. (2020), with 245 languages, as the study retested, and identifies Blasi et al. (2016) as having the largest language sample. The correction names the survey without adding a concept count: the paper uses 340 in its introduction and 344 later. - MINOR D4: "Most give away nothing"; "Most of the forty give away nothing at all"; "Almost all" comes from seven and "most of the rest" give nothing away
public/strata/the-word-that-gives-itself-away/data.json:perConcept; research/the-word-that-gives-itself-away/recheck-2026-09-29.mjs : Recount: 12 of 40 have accuracy at or below 0.025. Sum n*(accuracy-0.025) over the top seven divided by that sum over all forty is 0.7606937493726428. The replacement says about three quarters of the improvement over guessing and twelve at or below chance. - MINOR D5: 145 board languages, "one from each family in the study"
research/the-word-that-gives-itself-away/export-board.mjs:48-56; data/dataset.json; public/strata/the-word-that-gives-itself-away/board.json : Recount: 145 board languages in 145 families; the study contains 371 families. There are 1,152 complete lists in exactly 145 families, which is the board's selection rule. - WRONG D6: "words for breasts and mountain are long"
research/the-word-that-gives-itself-away/data/dataset.json; asjpcode.mjs:flatSymbols; recheck-2026-09-29.mjs : Mean symbol length per language is 4.0775706554 for breasts against 4.1716145826 averaged across the forty concept means. Knee is 5.3657458564 and mountain is 4.6259611036. "Knee" replaces "breasts"; no replacement figure is invented. - MINOR D7: without dropping duplicate Mandarin lists, "the model would meet a dialect of its own test language in training"
research/the-word-that-gives-itself-away/run.mjs:133-151; build.mjs:perGlotto; https://raw.githubusercontent.com/lexibank/asjp/v21/cldf/languages.csv : Direct CSV recount: 184 rows have Glottocode mand1415, all Family Sino-Tibetan. Family holdout excludes those together; the cheat arm includes the test language anyway. Choosing one list changes a language's weight in training and testing, not the family exclusion rule. - MINOR D8: the free shuffle collapses to "exactly one in forty"
public/strata/the-word-that-gives-itself-away/data.json:ladder.nullFree; recheck-2026-09-29.mjs : The measured accuracy is 0.024935053660035126, close to but not exactly 0.025. The metadata now says "about"; the dek distinguishes measurement from the arithmetic baseline. - MINOR D9: the film caption says "2.4933 per cent"
research/the-word-that-gives-itself-away/data/findings.json:ladder.nullFree; film/timeline.mjs:100-102 : 100 times the recorded accuracy is 2.4935053660035127, which rounds to 2.4935%. The encoded-audio check's old 2.4933% was a diagnostic literal, not its numeric assertion. That diagnostic now formats the actual rung and chance values; the assertion still uses the computed beat. - MINOR D10: "Every pitch in it is one of the numbers on this page"
research/the-word-that-gives-itself-away/film/make-audio.mjs:75-87; film/timeline.mjs:32-35; film/film.html : In addition to mapped accuracies, the code adds TONIC/2 = 55 Hz to the drone and plucks at TONIC*2 = 220 Hz. The caption, VideoObject, closing film text, film documentation and comments now describe those exceptions. - MINOR D11: "Every number here is produced under one rule: the model never sees the test language" or its family
research/the-word-that-gives-itself-away/run.mjs:122,133-151 : The headline satisfies that rule; the cheat arm uses all languages. The claim is narrowed to the headline in page prose and the film's rule text, rather than changing any training computation. - MINOR D12: both shuffled nulls remove all sound-meaning links and an above-chance result there would invalidate the pipeline
research/the-word-that-gives-itself-away/run.mjs:58-90; public/strata/the-word-that-gives-itself-away/data.json:ladder : The equal-length shuffle preserves the length-meaning link and scores 3.087872898843989%. The free shuffle scores 2.4935053660035127% against 2.5%. The page now distinguishes the two nulls and says a systematic excess in the free shuffle would expose an artefact. - MINOR D13: WHB printed -0.05; "Running the same comparison here" gives -0.047
https://pdfs.semanticscholar.org/e907/57eb47e9d40431821dae1a72d03537d8de4f.pdf, Table 4 and section 5; research/the-word-that-gives-itself-away/recheck-2026-09-29.mjs : The PDF prints -0.05 but does not name the correlation method there. Its Table 4 D/S values match the transcription. Recomputed Pearson(D,S) = -0.04581265858719753, rounding to -0.05; Spearman(D,S) = -0.08453743815571955; Spearman(accuracy,S) = -0.04738224891080841. The replacement states the numeric match to Pearson, rather than attributing an unprinted method name to the authors, and identifies this page's Spearman statistic. - CONFIRMED D14: "What did survive, in his words, were the pronominal nasals and the lateral in the word for tongue"
https://arxiv.org/html/2512.07543v2, section 4.3 and conclusion : Re-read: the Holman-40/Swadesh-100 survivors include I with nasal and stop features and thou with nasal; the discussion identifies a lateral/voiced effect for tongue among the strongest candidates. The adjudicated "not false" paraphrase is unchanged. - MINOR D15: "Every number above is recomputed from the committed artefacts"
research/the-word-that-gives-itself-away/analyse.mjs:259-269; public/strata/the-word-that-gives-itself-away/page.js:398; film/check-encoded.mjs : The release totals are explicitly transcribed in analyse.mjs; 184 is typed in page.js (the CSV recount confirms it); the old film-caption decimal was typed and wrong. The replacement retains the supported statement that the checklist is generated from the committed artefacts. - MINOR D16: "ASJPcode throws away ... nasalisation"
https://pdfs.semanticscholar.org/e907/57eb47e9d40431821dae1a72d03537d8de4f.pdf, paragraph after Table 2; research/the-word-that-gives-itself-away/asjpcode.mjs:8-13,98-104 : The PDF says ASJP has symbols for glottalization and nasalization. The parser documents * as nasalisation, and flatSymbols('m*o*ni') returns m,o,n,i. It is the model's flat reading that drops the mark. - MINOR D17: "Shared ancestry as the enemy ... There it misleads a method" in the more-data-wrong-tree relates note
public/strata/more-data-wrong-tree/index.html:292-296 (read only) : That page describes coincidental shared bases being read as ancestry by parsimony. The note now distinguishes that mistake from the ancestry confound controlled in this layer. The linked layer was not edited. - MINOR D18: Blasi's method and this one "have almost nothing in common except the database family" and agreement cannot arise from shared method
https://pure.mpg.de/rest/items/item_2344242_2/component/file_2344241/content, association method and filters; research/the-word-that-gives-itself-away/associations.mjs:14-17,25-33,60-76 : Re-read both: lineage balancing, within-language baselines, word-length controls and macroarea checks are shared, though the statistics, code and releases differ. The unsupported independence assertion is replaced with that distinction. - MINOR D19: "67 of 67 published sound-meaning signals come back with the same sign"
public/strata/the-word-that-gives-itself-away/data.json:keys.blasi2016; https://pure.mpg.de/rest/items/item_2344242_2/component/file_2344241/content, Results/Table 1 : The paper reports 74 published associations. The committed key has published = 74, onCore40 = 67, signAgrees = 67. The search description now specifies the 67 of 74 that concern these meanings.
What was done
- [fixed] Found by the codex fix-checker (D15): page.js's header comment still said no number in the file is typed by hand, while the Mandarin count 184 is a literal; the comment now says the model's results come from data.json and board.json and that a few source counts are transcribed.
- [fixed] Found by the codex fix-checker (D10): the film's entry in src/data/films.json (served at /tv/films.json) still carried the old "every pitch" description and the old film's size and hash; regenerated by the directing instance with scripts/tv-films-manifest.mjs from the corrected VideoObject and the rebuilt film.
- [declined] Found by the codex fix-checker (D1): research/the-word-that-gives-itself-away/run.mjs still labels its cheat arm "CHEAT: test family left in the training set", and results.json records that label. Left as it is on purpose: the handover's rule is to make the prose say what the code did, not to change the code, and the label is part of the committed run's record; the page and data.json now describe the arm correctly (it also trains on the test language).
- [fixed] D1: Corrected the dek, load-time ladder explanation, and labels authored in analyse.mjs; regenerated findings.json and public data.json, including the arms-table label. The film uses the regenerated ladder label. The experimental run.mjs, tune.mjs and historical results.json remain unchanged; presentation labels describe what those runs did.
- [fixed] D2: The table description and heading now name the final protocol and explicitly say that the 36-variant tune.mjs sweep is not shown.
- [fixed] D3: Named Erben Johansson et al.'s 2020 survey in the page, footer and research README, and scoped "published associations" to that survey.
- [fixed] D4: Corrected the OG description, visible introduction and markdown plain text; regenerated the plain meta from its frontmatter source.
- [fixed] D5: The load-time map caption limits its families to those with a complete forty-word list.
- [fixed] D6: Replaced breasts with knee in the long-word example.
- [fixed] D7: The load-time filtering explanation says deduplication prevents extra language weight; family holdout already excludes all the lists together.
- [fixed] D8: Both description copies now say "about one in forty", and the dek calls 1 in 40 the baseline against which 2.49% was measured.
- [fixed] D9: Corrected the caption and film documentation to 2.4935%; replaced the audio verifier's incorrect diagnostic literal with the data-derived decimal. No measurement threshold was changed.
- [fixed] D10: Qualified the film pitch claim in all authored copies, including the film's closing text, and named the extra drone and pluck pitches.
- [fixed] D11: Limited the general exclusion rule to the headline in the page and film.
- [fixed] D12: Distinguished the retained length link from the free-shuffle null.
- [fixed] D13: The dek and load-time WHB discussion now distinguish the statistics and identify the Pearson calculation as a match to the printed value.
- [fixed] D15: Limited the blanket recomputation promise to the generated checklist.
- [fixed] D16: Separated ASJP's nasalisation notation from its removal by this model.
- [fixed] D17: Corrected the relates note in frontmatter and regenerated its nearby-card teaser using the existing generator in an isolated scratch tree containing only this layer's HTML.
- [fixed] D18: Replaced the claim of methodological independence with the shared and differing features established by the sources.
- [fixed] D19: The search description now says all 67 of the 74 published signals concerning these meanings match in sign.