The Groundtruth Seam · recommendations tested later

Did the Error Bars Hold?

First reproduce a published calibration result. Only then test how often matchable historical recommended-value intervals contain a later recommendation, with every matching and exclusion rule left where you can move it.

The published anchor, first

Make the old result come back

Henrion and Fischhoff, 1986checking

Their Table I reports two benchmarks at the same threshold. A calibrated normal distribution puts 2% outside 2.33 standard uncertainties. The browser recomputes that tail probability below. The table also reports 57% among 40 historical recommended values, and the article does not print those 40 rows. It does name them: five constants across the eight reviews of 1929 to 1969, with residuals taken against the 1973 recommendations. We transcribed the forty values and uncertainties by hand from those primary reviews and committed them beside this page, and the browser recounts them live. The reconstructed count is loading.

Two other statistics in the same Table I row do not come back. The committed rows give a root-mean-square normalized residual of pending against their published Birge ratio of 7.42, and an interquartile index of pending against their published 22%. The article prints neither its 40-row working table nor its atomic-mass-scale transformations, so the difference cannot be located from what was published. This reconstruction takes each value and uncertainty exactly as its review printed it, with no probable-error conversion and no mass-scale transformation.

The two thresholds are not the same instrument. This anchor counts recommended values that fall outside 2.33 quoted standard uncertainties. Every measurement below counts historical intervals of one standard uncertainty by default. Reproducing the anchor tests the pipeline, and it does not produce the same statistic as this page's own headline.

published normal benchmark
2%
recomputed normal tail
pending
published historical index
57% of N = 40
recomputed from the 40 committed rows
pending

Layer 1 · the modern record

Now move the interval

We could not find this published as of 2026-08-01, having searched NIST CODATA background, archive, bibliography and version-history pages; CODATA adjustment papers; PDG publications, API documentation and previous-edition pages; Henrion and Fischhoff 1986; Bailey 2017 and citing literature; and exact-phrase web searches for historical CODATA coverage, PDG interval coverage and surprise index.

Modern coverage benchloading the ledger
contained
pending
coverage
pending
signed median miss
pending

The row ledger is loading.

Primary records: NIST CODATA version history and ASCII archive and the official PDG historical database. Prior art includes Henrion and Fischhoff 1986, Bailey 2017, and the representative 2014 to 2018 comparison in the 2018 CODATA report.

Layer 2 · denominator under pressure

The headline may be made by the matching rule

The sophisticated dismissal: today's value is not truth, correlated derived constants repeat the same information, unchanged recommendations make easy successes, and the 2019 SI redefinition makes some comparisons definitional.
Sensitivity benchready
contained
pending
coverage
pending
rows removed
pending

computing quantity-cluster bootstrap

Inspect the largest signed residuals under these choices
quantityeditionsigned missstate

The check

The browser recomputes every rate from the shipped row ledger. Choices: historical point estimate or interval overlap; latest completed CODATA adjustment 2022; latest PDG comparison edition 2026 while historical rows stop at 2025; asymmetric errors use the side facing the comparison value; one sigma means the quoted standard uncertainty; boundary equality counts as contained.

published anchorpending
anchor reconstruction40 rows, five constants across the eight reviews of 1929 to 1969, hand transcribed from the primary papers because Henrion and Fischhoff print aggregates only; residuals taken against the 1973 CODATA recommendations; values and uncertainties exactly as printed, no probable-error conversion, no atomic-mass-scale transformation; committed as research/error-bar-coverage/data/anchor-rows.json and shipped byte for byte in the page
modern defaultpending
data islandpending
CODATA selection ledger1,795 finite historical nonexact rows; 477 lack an exact display-name and unit match in 2022; 1,318 retained, including 191 rows whose 2022 target is exact under the revised SI
PDG selection ledgermass and lifetime only; 10,337 eligible row instances; 224 rows in ambiguous duplicate edition pairs dropped; among 9,702 unique historical rows, 285 lack a unique 2026 row and 172 change unit; 9,245 retained
matching conventionCODATA display name and unit must match exactly; PDG identifier and unit must match exactly; no manual renames
latest uncertaintyzero for exact 2022 SI definitions; otherwise the published 2022 CODATA or 2026 PDG uncertainty
free choicessigma width, point or overlap metric, current or next target, all or changed rows, SI filter, earliest-row collapse, 400 seeded quantity-cluster bootstrap replicates

Run node research/error-bar-coverage/verify-error-bar-coverage.mjs.

The open edge

The latest recommendation is still a measurement

This page does not infer misconduct, confirmation bias or a causal bandwagon effect. It audits recommended values, not every input experiment. Rows are correlated, survival into the latest table is selective, and today's recommendation can move. A prospective rerun after CODATA 2026, or a reconstruction using full input covariance matrices, remains undone.