The Groundtruth Seam · recommendations tested later
Did the Error Bars Hold?
First reproduce a published calibration result. Only then test how often matchable historical recommended-value intervals contain a later recommendation, with every matching and exclusion rule left where you can move it.
Make the old result come back
Their Table I reports two benchmarks at the same threshold. A calibrated normal distribution puts 2% outside 2.33 standard uncertainties. The browser recomputes that tail probability below. The table also reports 57% among 40 historical recommended values, and the article does not print those 40 rows. It does name them: five constants across the eight reviews of 1929 to 1969, with residuals taken against the 1973 recommendations. We transcribed the forty values and uncertainties by hand from those primary reviews and committed them beside this page, and the browser recounts them live. The reconstructed count is loading.
Two other statistics in the same Table I row do not come back. The committed rows give a root-mean-square normalized residual of pending against their published Birge ratio of 7.42, and an interquartile index of pending against their published 22%. The article prints neither its 40-row working table nor its atomic-mass-scale transformations, so the difference cannot be located from what was published. This reconstruction takes each value and uncertainty exactly as its review printed it, with no probable-error conversion and no mass-scale transformation.
The two thresholds are not the same instrument. This anchor counts recommended values that fall outside 2.33 quoted standard uncertainties. Every measurement below counts historical intervals of one standard uncertainty by default. Reproducing the anchor tests the pipeline, and it does not produce the same statistic as this page's own headline.
Now move the interval
We could not find this published as of 2026-08-01, having searched NIST CODATA background, archive, bibliography and version-history pages; CODATA adjustment papers; PDG publications, API documentation and previous-edition pages; Henrion and Fischhoff 1986; Bailey 2017 and citing literature; and exact-phrase web searches for historical CODATA coverage, PDG interval coverage and surprise index.
The row ledger is loading.
Primary records: NIST CODATA version history and ASCII archive and the official PDG historical database. Prior art includes Henrion and Fischhoff 1986, Bailey 2017, and the representative 2014 to 2018 comparison in the 2018 CODATA report.
The headline may be made by the matching rule
computing quantity-cluster bootstrap
Inspect the largest signed residuals under these choices
| quantity | edition | signed miss | state |
|---|
The check
The browser recomputes every rate from the shipped row ledger. Choices: historical point estimate or interval overlap; latest completed CODATA adjustment 2022; latest PDG comparison edition 2026 while historical rows stop at 2025; asymmetric errors use the side facing the comparison value; one sigma means the quoted standard uncertainty; boundary equality counts as contained.
| published anchor | pending |
|---|---|
| anchor reconstruction | 40 rows, five constants across the eight reviews of 1929 to 1969, hand transcribed from the primary papers because Henrion and Fischhoff print aggregates only; residuals taken against the 1973 CODATA recommendations; values and uncertainties exactly as printed, no probable-error conversion, no atomic-mass-scale transformation; committed as research/error-bar-coverage/data/anchor-rows.json and shipped byte for byte in the page |
| modern default | pending |
| data island | pending |
| CODATA selection ledger | 1,795 finite historical nonexact rows; 477 lack an exact display-name and unit match in 2022; 1,318 retained, including 191 rows whose 2022 target is exact under the revised SI |
| PDG selection ledger | mass and lifetime only; 10,337 eligible row instances; 224 rows in ambiguous duplicate edition pairs dropped; among 9,702 unique historical rows, 285 lack a unique 2026 row and 172 change unit; 9,245 retained |
| matching convention | CODATA display name and unit must match exactly; PDG identifier and unit must match exactly; no manual renames |
| latest uncertainty | zero for exact 2022 SI definitions; otherwise the published 2022 CODATA or 2026 PDG uncertainty |
| free choices | sigma width, point or overlap metric, current or next target, all or changed rows, SI filter, earliest-row collapse, 400 seeded quantity-cluster bootstrap replicates |
- CODATA rows are algebraically correlated. PDG rows repeat recommendations across editions. No raw-row rate is given an independent-binomial confidence interval.
- The 2019 SI redefinition fixed several defining constants. Those rows are visible in the raw descriptive rate and removable in the sensitivity bench.
- Limits, ranges, confidence-level bounds, zero uncertainties, ambiguous PDG duplicates and PDG properties other than mass or lifetime are outside the denominator.
- The source tables use rounded display values. Equality at the displayed boundary is counted as contained; no claim is made about hidden digits.
Run node research/error-bar-coverage/verify-error-bar-coverage.mjs.
The latest recommendation is still a measurement
This page does not infer misconduct, confirmation bias or a causal bandwagon effect. It audits recommended values, not every input experiment. Rows are correlated, survival into the latest table is selective, and today's recommendation can move. A prospective rerun after CODATA 2026, or a reconstruction using full input covariance matrices, remains undone.