The Groundtruth Seam · what an evaluator does when the data will not agree
The Half-Lives That Would Not Agree
NUBASE2020 flags a contested half-life by printing a Birge ratio beside it. Count the flags and the answer depends on the matching rule: half-life comments carry a numeric Birge quantity, use the literal words "Birge ratio". Rebuild gold-196's published 6.165(11) d from three disagreeing measurements, then audit every flagged ratio against the numbers printed next to it.
A half-life looks like a fact. It is usually an average, and sometimes the things being averaged do not fit inside each other's error bars. When that happens the evaluators of NUBASE2020 say so in the margin of Table I, printing a single number that measures how much wider the scatter is than the quoted uncertainties allow. This page starts with the worst case they printed.
Three laboratories, one gold nucleus
Gold-196 has been measured three times in ways NUBASE2020 kept. The intervals below are one standard deviation wide, and no two of them overlap. Switch any of them off and every number underneath is recomputed from the survivors.
The dashed line is the inverse-variance weighted mean. The solid gold line is the plain arithmetic mean. Which one NUBASE publishes depends on the number in the middle panel.
With all three switched on, B is above 4. NUBASE2020's own text says what happens then. Quoting the paper's section 3.2: In rare cases when χn is larger than 4, all individual uncertainties are considered to be irrelevant and the arithmetic (unweighted) average is adopted That is the sentence the bench is executing. The precision-weighted answer, which leans hardest on the most confident laboratory, is thrown away in favour of one that treats all three equally.
Does the instrument agree with the published record?
NUBASE2020 prints two numbers for ground-state gold-196: a recommended half-life in its machine-readable table, and a Birge ratio in its Table I comment. Both are reproduced below from the three input measurements alone, live in this page, before any number of ours appears. Set the toggles above back to all three to see the comparison the evaluators made.
Anchor reproduction, gold-196 ground state
| quantity | published | recomputed here | verdict |
|---|
The published row of the fixed-width NUBASE2020 table reads loading, where the half-life field is . The comment that supplies the inputs reads, verbatim: loading.
The three inputs are Hirose, Kikunaga and Ohtsuki (2011), Lindenberg and seven coauthors (2001), and Ikegami, Sugiyama, Yamazaki and Sakai (1963). The 2011 and 2001 values sit combined standard deviations apart. Nothing about the arithmetic decides which is right. The evaluators' rule decides only that neither is trusted enough to dominate.
Now make the dismissal fail
Pick any flagged nuclide. Every estimator below is rebuilt from that entry's printed inputs at the current settings, and the counters underneath are rebuilt across all entries whose inputs the comment prints in full.
| estimator | central value | uncertainty | relation to NUBASE's rule |
|---|
What moves across the whole flagged set
Recomputed live over every entry whose inputs are printed, at the current scale and thresholds. The published NUBASE settings are scale 1.00, thresholds 2.50 and 4.00.
Two of these controls are vacuous on purpose, and the page will not pretend otherwise. Multiplying every quoted σ by a common factor s divides every Birge ratio by exactly s and leaves both the weighted mean and the unweighted mean untouched, by arithmetic. The button that drives B to 1 therefore always succeeds at s = B, and proves nothing. What is not vacuous is the counter labelled recommended value moved: it asks how many published half-lives would change their central value, not merely their error bar, if the evaluators had drawn the line somewhere else. At NUBASE's own settings that counter reads , because it is measured against those settings. Drag threshold 2 down toward 2.5 and watch it climb.
The leave-one-out sweep answers the other half of the dismissal directly, and it does not answer in our favour. For every entry with three or more inputs it drops the measurement contributing most to chi-squared and recomputes. On the flagged set at NUBASE's own thresholds, . So the dismissal is often correct: one laboratory frequently is the disagreement.
Gold-196 is one of those cases. Drop and its ratio falls from to , which is inside the band where the weighted mean survives untouched. Iridium-194 is the opposite case: dropping its worst input moves the ratio only from to , because all three of its measurements disagree with each other. Load either into the bench and switch inputs off to watch it happen. What the flag records is not a diagnosis. It is a refusal to pretend the inputs are compatible, and it says nothing about which one is wrong.
The audit: does each printed ratio follow from the numbers printed beside it?
This is the part nobody has to take on trust. Each row recomputes B from that comment's own printed inputs, under NUBASE's stated convention, and compares it with the ratio the evaluators printed. Rows are shaded by verdict. Click a row to load it into the bench above.
| nuclide | N | printed B | recomputed B | gap | verdict | NUBASE value | load |
|---|
Forty-six flags, and how many survive their own arithmetic
The matching rule comes first, because the headline count depends on it. A NUBASE2020 Table I comment is a block of consecutive lines sharing a nuclide, a state, and a property code. Restricting to blocks whose property code is T, the half-life property, and whose text contains the string Birge:
| matching rule | count | what it excludes |
|---|
The gap between and is a single comment. Gallium-76 prints Birge B=2.7 rather than Birge ratio=2.7, and B is the paper's own symbol for the quantity. A literal-string search reports . A search that reads the notation reports . We report both and take as the scientific answer.
On that set, and stated in full: we could not find this published as of 2026-08-01. Of the NUBASE2020 half-life comments carrying a numeric Birge quantity, print their complete input list. Recomputing each from those inputs under the paper's own equation 4, reproduce the printed ratio to the precision at which it was printed, more land within one percent, do not reproduce at all, and cannot be checked because its inputs are not printed. The largest printed ratio in the set is for ground-state gold-196.
The three that do not reproduce are named on the page, not buried: . Their gaps are not small. Tin-134's comment prints beside inputs that give ; rhodium-101m prints beside inputs that give ; tin-137 prints beside inputs that give .
For tin-134 there is a reconstruction, and it is ours rather than the source's. Its comment prints 15Lo04=0.89(0.2). Reading that uncertainty as 0.02 s instead of 0.2 s, and changing nothing else, gives , which prints as . We cannot tell from the published article whether the printed 0.2 is a typographical slip or whether the evaluators averaged something else, so we report the disagreement and the candidate reading separately. For rhodium-101m and tin-137 we have no reconstruction at all. Rhodium-101m carries a further internal tension worth naming: its published uncertainty of equals the internal uncertainty of exactly those three printed inputs, which under the paper's own rule is what you adopt when B is at most 2.5, not when B is 3.22.
We searched for a prior census or audit of these comments in the following specific places: the full NUBASE2020 article at DOI 10.1088/1674-1137/abddae including all 180 pages of Table I and its reference list; the NUBASE2016 predecessor, which prints no gold-196 Birge comment at all; the NNDC Nuclear Science References database, by keynumber for the three gold-196 inputs and by topic for NUBASE and Birge ratio; and open web searches for the phrases NUBASE2020 Birge ratio census, worst Birge ratio NUBASE half-life, and NUBASE2020 10.86 196Au. Every component ratio in this audit is itself published, in the source, by its evaluators. What we could not find published is the enumeration, the recomputation of each ratio from its own printed inputs, or the resulting verdict. If it exists in a conference slide, an evaluator's notebook, or an unindexed report, this claim is wrong and the reproduction above still stands.
The check
Everything above is computed in your browser from a -entry JSON record embedded in this page, which is the same file the offline verifier reads. No number on this page is a literal typed under a claim that it was computed. The free choices are these, and each one changes an answer:
| source edition | NUBASE2020, literature cutoff 2020-10-30. NUBASE2016 gives a different gold-196 recommendation and no Birge comment for it. |
|---|---|
| Birge convention | B = sqrt(χ2/(N−1)) about the inverse-variance weighted mean, the paper's equation 4. Dividing by N instead gives for gold-196; calling reduced chi-squared itself the ratio gives ; taking residuals about the adopted unweighted mean gives . |
| uncertainty meaning | Every quoted uncertainty is treated as one standard deviation, which is NUBASE2020's stated policy. Statistical and systematic parts are combined in quadrature, also its stated policy. Doubling every σ halves every B and leaves both means unchanged. |
| parenthetical convention | A parenthetical containing a decimal point is absolute (6.183(0.010) is ±0.010). A bare integer applies to the trailing digits (32.643(260) is ±0.260, 71.25(42) is ±0.42). |
| asymmetric uncertainties | Symmetrised by NUBASE2020's own equations 12 and 9, which shift the centre as well as the width: m = X + 0.64(a−b), σ2 = (1−2/π)(a−b)2 + ab. Two entries use this. Using the naive midpoint instead moves thorium-217 from agreeing to disagreeing. |
| comment grouping | Consecutive lines sharing nuclide, state and property code are one comment. A nuclide's inputs are routinely split across a line break, and gold-196's are: line carries two inputs, line carries the third and the ratio. |
| state discipline | Ground states and isomers are separate rows. of the flagged entries are isomeric or higher states, marked m, n or p, and must never be merged with their ground state. |
| which inputs count | Values printed after "others", "other", "supersedes", or "discrepant" are excluded from the average, as the comment says. This is a reading of English prose and is the softest link in the chain. Every comment's full text is shown beside its entry so you can disagree. |
| verdict thresholds | "Exact" means the gap is at most half a unit in the last decimal place the source printed. "Near" means within one percent. Both are our choices, not the source's. |
| units | The Birge ratio is dimensionless and invariant under any common linear unit change, so no conversion is applied. Two entries (helium-5 and lithium-10 n) compare resonance widths in keV rather than times, and are labelled. |
| independence | No covariance is published for any of these input sets, so every calculation here assumes uncorrelated measurements. Shared systematics between laboratories would change every ratio and cannot be recovered from the comment. |
The primary records are committed beside the verifier: the fixed-width evaluated table nubase_4.mas20.txt as fetched from the IAEA mirror, the verbatim comment lines nubase2020-birge-comments.txt with their line numbers in the article's text layer, and the distilled record birge-halflife-census.json that this page embeds.
Run the independent checker:
node research/halflife-disagreement/verify-halflife-disagreement.mjs
It parses the verbatim comment lines with its own parser, recomputes every ratio with its own arithmetic, then extracts this page's inline engine, executes it, and asserts the two agree. It finishes by deliberately corrupting inputs and conventions and asserting the checks turn red.
This is a census of flags, not of disagreements
The honest limit is blunt. This page counts the half-lives whose evaluators chose to print a Birge ratio in a comment. It is not a census of every genuinely contested half-life, and the two sets are not the same size. A nuclide measured once has no ratio to print. A nuclide whose discrepant measurement was rejected outright before averaging may carry no ratio either, because the surviving values agree. And a comment can print a ratio for measurements the evaluators went on to discard, as mercury-197's does: its 6.7 is explicitly what the ratio would be if a strongly conflicting 1966 value were included, and it was not included.
So the frequency of these flags says nothing about how often nuclear half-lives disagree. Treating as a prevalence would make the central claim false. The number is a property of an editorial practice: how often five evaluators, reading the literature up to 2020-10-30, decided a disagreement was worth naming in the margin.
The next honest step is one this page does not take. A real census of contested half-lives would need every accepted and rejected measurement extracted from ENSDF and the DDEP evaluations as well as NUBASE, joined to the evaluator decisions that excluded them, and carried forward past the 2020 cutoff into the primary literature since. That work needs the rejected values, which are exactly the ones an evaluated table is designed not to carry.
Primary source: F. G. Kondev, M. Wang, W. J. Huang, S. Naimi and G. Audi, "The NUBASE2020 evaluation of nuclear physics properties", Chinese Physics C 45 (2021) 030001, doi:10.1088/1674-1137/abddae, Creative Commons Attribution 3.0. The statistic is Raymond T. Birge, "The Calculation of Errors by the Method of Least Squares", Physical Review 40 (1932) 207, doi:10.1103/PhysRev.40.207.