Arm B, the living
Fifteen words. One pass.
Press start. A fixed vocabulary appears in one of 60 sealed orders, one word every 600 milliseconds. Then type whatever remains. Scoring is exact and local. Your typed words never leave this browser.
The 60 orders are counterbalanced: four Latin squares, so every word serves every position in exactly four of them, and word difficulty cannot pose as a position effect. Only the first list you complete in a visit can be sent; later lists are practice and stay on this device.
The volunteer counter
0 arrivals, never a finished sample.
Reading the arm.
0
Descriptive mean across these one-trial volunteers.
Bankroll against no contrast: 1.00x now, 1.00x at its peak
Bankroll against no contrast: 1.00x now, 1.00x at its peak
| position | recalled | rate | 99.667% confidence sequence |
|---|
Why these intervals: a confidence sequence is valid at every moment simultaneously, including after every visitor refreshes this page. A conventional 1.96-standard-error interval assumes a fixed stopping point. Reusing it under a live counter would quietly spend more false-positive risk at every look.
The 15 position sequences divide a 5% familywise error budget by 15. The two preregistered participant contrasts divide another 5% familywise budget by two, so each contrast's e-process pays alpha 2.5% and its threshold is forty-to-one. The bankroll shows two numbers on purpose: the standing wealth, which can fall back when evidence reverses, and the peak, which is what the threshold rule reads because a Ville crossing is permanent evidence. A peak is never passed off here as the current e-value. Volunteer arrival is not random sampling, so none of this makes the crowd representative.
Arm A, the published record
The table passes. The reproduction gate stays closed.
Murdock reported six between-group conditions, 103 introductory-psychology students in total, and 80 lists per participant across four sessions. The suffix in each condition is seconds per word. The page recomputes the participant total and checks every transcribed aggregate cell against a separately stored printed check.
| condition | N | transcribed mean | printed check | transcribed SD | agreement |
|---|---|---|---|---|---|
| sum / weighted arithmetic | computing | computing | weighted mean is derived here, not printed by the paper | ||
Checking aggregate cells.
Checking the derived total-time regularity.
Comparison gate: checking
The public Penn archive contains 1,200 usable rows per condition, consistent with 15 people doing 80 trials, while the paper reports condition Ns from 15 to 19. The 10-word file carries a 1,201st row holding only the malformed value 0. The README documents 88 as an extra-list intrusion but does not explain observed out-of-range codes 0, 16, 31, 41, and 50. No explicit redistribution license was found. This page therefore ships neither those rows nor a curve derived from them.
Arm C, the machines
The same report format, not the same memory.
A fresh model context receives the same 15 nouns as visible text and is asked to report them, once for each of the 60 counterbalanced orders the living arm deals, so word identity and position are unconfounded for machines too. Every studied token remains available in the model's active context. That is in-context extraction, not human encoding followed by recall. A similar edge shape would not identify a shared mechanism.
Looking for supervisor-supplied raw runs.
0
generated
0
0 refusals
0 unparseable
0 rail errors, no completion returned
Census roster: unknown
Sampling parameters as recorded on each draw: unrecorded
| model | draws | parsed | refusals | unparseable | rail errors | mean words reported |
|---|
| position | reported | rate |
|---|
The parser applies Unicode NFKC and lowercase, splits only on comma, newline, or semicolon, trims surrounding ASCII punctuation, then exact-matches whole tokens against the order named by each run's protocol id, so a word scores at the position it was shown, not at a canonical position. It gives no fuzzy, stem, semantic, or manual credit. Each raw draw becomes value, refusal, unparseable, or rail error; completions all stay in the position denominator, and rail errors are counted in the open rather than hidden inside it.
The sealed instrument
The analysis had to exist before the crowd.
The literal below is part of this HTML, not recalculated to flatter the current file.
3631e768af9fd774b1a61c0c48a0a3dbc8f6d3f53f01d11e35027075859d1796
Checking the fetched file against the printed hash.
The check
What this page can fail to know.
Sources and audit trail
The record this page leans on.
Murdock, 1962
The primary paper for Table 1, the six conditions, 103 participants, 80-list procedure, and its fitted descriptions.
Penn archive
A public trial archive audited by the scout. It is not redistributed here because completeness and licensing remain unresolved.
Liu et al., 2024
Long-context question answering and key-value retrieval often lost information in the middle. It did not test free recall.
Guo and Vosoughi, 2025
Position effects across language-model classification labels and reordered summarization. Machine serial-position effects are not a novelty claim here.