The Groundtruth Seam · a portal across seven layers

Every Plus-or-Minus Is a Decision

Six decisions stand between a set of measurements and the sentence these agree, and every one of them is taken before anybody reads the numbers. Here they are, on real data from five layers of this archive, with the answer recomputed live as you move them. The opening case is a row that appears twice in this archive with two different signs after it.

One row, two documents, two verdicts

In 1963 a committee recommended a value for the charge on the electron. Two layers of this archive carry that recommendation. They agree on the number, 4.80298 in units of ten to the minus ten electrostatic units, and they print different things after the sign, because they read different pieces of paper.

Did the Error Bars Hold? transcribed it from Cohen and DuMond's 1965 review in Reviews of Modern Physics, where it is a standard deviation, and ships it as 4.80298 ± 0.00006. The Charge That Crept read the contemporaneous sheet from the National Bureau of Standards, which prints a wider figure and says of it, in its own words, Based on 3 std. dev., and ships it as 4.80298 ± 0.00020 tagged as a three-sigma limit.

Neither layer is wrong about its own document, and this page corrects neither. Two of the three documents involved are reachable and were opened; the third, Cohen and DuMond's 1965 review itself, returned 403 on every route tried, so the sentence above about what that review prints rests on the anchor layer's own transcription of it and not on a reading here. The card is NBS Misc. Publ. 253, and its uncertainty column is headed Est. error limit over a footnote reading, in full, Based on 3 std. dev., applies to last digits in preceding col. The adjustment's own uncertainty is printed in the comparison table of the 1973 adjustment as 1.60210(2) in units of ten to the minus nineteen coulombs, 12 ppm, which in the units of the other layer is , and rounds to the 0.00006 it ships.

So the two signs are one sigma and a policy. Taking the adjustment's sigma as its printed 12 ppm, the card's limit is times it rather than exactly three; taking instead the 0.00006 the anchor layer ships, it is 3.33. Both inputs are themselves rounded, and the card's limits are rounded up in turn, so the honest statement is that the printed limit is about three and a half times the adjustment's standard deviation, and that the exact multiple cannot be recovered from figures printed to two significant digits. The coulomb column of the same card shows the same thing and not more: a sigma of 2 in the last digits against a printed limit of 7, a ratio of 3.5 that three-times-then-round-up cannot produce on its own, since three twos are six.

Then ask the only question anybody actually wants answered. The elementary charge has been exact since the 2019 revision of the SI, so we know what the committee was aiming at. Did they miss?

The same row, read three ways

The sign, as some document prints itimplied sigmadistanceverdict
computing

Distance from the exact modern value loading, in units of the sigma each reading implies. The central value never moves.

The committee is either three and three quarters of its own error away from the truth, or one and a tenth, depending on which of two honest layers of this archive you happened to open, and the entire difference is a presentation decision taken in 1963 and announced in a five-word footnote that then travelled separately from the number. That is the whole of this page in one row: the number is not in dispute, and the verdict is.

The same bureau, on the same kind of card, then published the factor of three twice more, once inside the number and once outside it. Its 1971 card for the following adjustment carries the footnote Based on 1 std. dev.; applies to last digits in preceding column, and immediately beneath it:

These values may be in conflict with data available since the Taylor, Parker, Langenberg review. Pending a complete new readjustment of the constants, it would be prudent to multiply the above uncertainties by 3. NBS Special Publication 344 (1971), footnote to the table of recommended values

A standards body telling its readers, in print, to triple its own published uncertainties. The 1963 card put the three inside the sign and said so in five words. The 1971 card put the sign at one sigma and put the three in a sentence next to it. Both are honest. Neither number means anything until you have read the footnote, and the footnote is the part that gets dropped when a value is copied into the next table.

It is also not the only place in this archive where a convention is recorded twice and differently. Did the Error Bars Hold? keeps its forty rows in two artefacts, a JSON data file and a standalone Python reproduction, and they disagree about the 1952 review: the JSON records its uncertainties as standard errors, the Python tags them as probable errors and converts them at 1.48. Same five numbers, two conventions, one directory. This one turns out to be harmless: at the published threshold both labellings give the same count, because none of the five 1952 rows sits near enough to the line for a factor of 1.48 to carry it across. Found, measured, and immaterial, which is worth saying with exactly the same care as the case where it decides everything.

The decisions

Six of them, and one that comes before all six. Each is a real choice, each has been taken differently by real published bodies, and each is visible somewhere in the seven layers this page walks.

  1. Is there a plus-or-minus at all? Not every uncertainty is a sigma. A rounded figure carries an interval whose endpoints are exact and whose width the printer chose. A table of integer counts with a category called inconclusive carries no uncertainty until somebody decides how to score that category. The Room Inside a Number · The Answer That Can't Be Wrong
  2. What does the sign mean? A probable error is 0.6745 sigma. A three-sigma error limit is three of them. Millikan's 1913 figure is neither: The Charge That Crept quotes his own paper giving it as estimated limits of uncertainty rather than the so-called probable errors, which relates it to a standard deviation not at all, and that layer therefore refuses to convert it. So does this one. The Charge That Crept, which carries five conventions in one column and converts between them never
  3. What scale are the numbers on? Twentieth-century atomic weights ran on three incompatible scales, and the Avogadro constant moves with them. The conversions are small, exact, and run in opposite directions. Did the Error Bars Hold?, whose eight Avogadro rows sit on all three
  4. Which rows are in? One layer here ships Birge's 1941 recommendation; another deliberately omits Birge's 1941 and 1944 rows because it could not verify them from a primary source. Same archive, two inclusion rules, both defensible. Did the Error Bars Hold? · The Charge That Crept
  5. Are they independent? The variance of a difference carries a cross term. Assume independence and you have assumed a number nobody measured. The Gap That Changed Sides · Two Rulers That Will Not Agree
  6. How are they combined, and may you refuse? Weighted, unweighted, median, inflated by the Birge ratio, scaled by the PDG factor, or thrown out entirely above a threshold. Refusal is a rule too, and it is the one this page found hardest to find an example of. The Half-Lives That Would Not Agree, which watches five evaluators refuse an average at a numeric threshold
  7. What do you compare against, and is it also a measurement? It always is, until the day somebody defines it into exactness. The Gap That Changed Sides, where the experiment held still and the theory moved

The board

Here are forty rows: five constants, as recommended by eight reviews between 1929 and 1969, hand-transcribed from the primary papers by Did the Error Bars Hold? and shipped beside that layer. In 1986 Henrion and Fischhoff reported that 57% of them fell more than 2.33 of their own standard uncertainties from the then-current value, where a calibrated set would put 2% outside. They never printed the rows. Their paper names the five constants and the eight reviews and gives the aggregate; the forty rows below are a hand reconstruction from the primary papers, which is why "does the 57 percent come back" is a question with an answer rather than a tautology. That number is the reason anybody says physicists underestimate their error bars.

Move the decisions and watch it. Every level below is a rule some published body has used or a rule a member of this spine actually applies.

surprise rate
rows outside
a calibrated set
times too many

computing

each of 144 rules   the rule you have set   what a calibrated set would give

Two things are true at once on that strip, and telling them apart is the only reason this page exists.

The number moves. Across the whole grid of 144 rules the surprise rate runs from to . At the published 2.33 threshold, of rules land within a point of the published 57%. Eight rules do, five of them landing on 57.50% and three on 57.14%, and the naive one, values and uncertainties exactly as printed, is among them; that is the rule the member layer says on its own face it used. So the reproduction is a fact about a convention, not a fact that was going to come back whatever you did.

The conclusion does not move. Every one of the 144 rules returns a rate above what a calibrated set would give at its own threshold, by a factor of at least . Whatever you decide about probable errors, mass scales, which reviews to admit and what to compare against, these intervals did not cover. That is worth having precisely because it survived the attempt to break it.

Which decision does the work? Averaging over everything else, the swing each one contributes, in percentage points of the surprise rate:

Main effect of each decision

decisionswinglevels, and the mean rate under each
computing

Set the threshold aside as the most obvious knob and the ordering is worth sitting with: which rows you admit moves the answer more than what you take the sign to mean, and both move it more than which epoch you measure against.

The reference is also a measurement

The forty rows are scored against the 1973 adjustment of the constants, which is what Henrion and Fischhoff used. The seventh decision says that a reference is a measurement until somebody defines it into exactness, so apply the instrument to the instrument. Three of those five constants have since been fixed exactly by the 2019 SI, and the other two are known far better than they were. What does the 1973 table look like under its own test?

The 1973 recommended values, scored by the 1986 rule

constant1973, one sigmanowdistanceverdict
computing

This is the same operation the anchor performs on its forty rows, one epoch later: a recommended value, divided by its own quoted uncertainty, against a value we hold better. The 1973 table plays the part the 1929 to 1969 tables played, and today plays the part 1973 played. The threshold is the anchor paper's own, and the modern values are fetched from the NIST table rather than recalled.

Four of the five are more than 2.33 of their own standard deviations from the value we now hold. It is tempting to stop there, and the first draft of this page did, and it was wrong. Those are not four failures. They are one failure, counted four times, and the paper that printed the values printed the reason two pages later.

Table 33.4, the same paper, page 722

paircorrelationpaircorrelation
computing
pairs above 0.9
independent quantities
condition number

Footnote a to the recommended-value tables says it in words, and this page quoted the footnote before it obeyed it: the uncertainties of these constants are correlated, and therefore the general law of error propagation must be used in calculating additional quantities requiring two or more of these constants.

The eigenvalues of that matrix say how many numbers are really there. Its participation ratio is of five. Two further things fall straight out of it, and both hold. The three constants that move up together, e, h and me, are correlated with one another between 0.904 and 0.991, and the one that moves the other way, NA, is anti-correlated with all three, so the sign pattern of the deviations is exactly the one a single shared shift would produce. And the constant that passes the test is alpha-1, whose largest correlation with any of the other four is , the only one of the five that is not mostly the same number again.

So what is the right test? The general law of error propagation gives one: the joint statistic z' R-1 z, which asks how surprising the whole set of deviations is given how the set is tied together. It is one line of arithmetic and this page will not print its answer, for a reason worth more than the answer.

That matrix cannot be inverted from the precision it was printed at. Every coefficient is given to three decimals, so each is known only to within five ten-thousandths. Take the box that describes and evaluate the joint statistic at all 1024 of its corners: it comes back between and , a factor of , and of those corners are not a valid correlation matrix at all. The number this page would print is a fact about the last digit that was printed, not about the constants. The condition number is , so that is exactly what a condition number that size means.

Which is decision zero and decision four arriving together, on the page that named them. The Room Inside a Number is the member that says a printed figure is an interval whose width the printer chose; here the printer chose a width that will not carry the calculation the same paper tells you to do. The honest output is the range, and the range is four orders of magnitude wide.

What survives. The deviations are real, they are coherent, and they run in the direction the published correlations require. The 1973 recommended values did move by more than their own quoted uncertainties, and one shared shift of roughly three of those sigmas is enough to produce all four. What does not survive is the count: four of five is one piece of evidence wearing four coats, and a page about decisions had no business reporting it as four until it had read the matrix.

None of which makes the 1973 paper careless. Read what it says about itself, in the footnote where it defines the very statistic this page has been using:

Recall that σI, the uncertainty determined by internal consistency, is the expected uncertainty in the mean as determined by the a priori uncertainties, σi, assigned each individual measurement; and that σE is the expected uncertainty as determined by how much each individual measurement deviates from the weighted mean in comparison with its a priori uncertainty σi. The Birge ratio, RB, is defined as σEI and is related to χ² by RB = [χ²/ν]1/2, where ν is the number of degrees of freedom. The expectation value of χ² is ν, and thus of RB², unity. RB>1 generally implies that either the σi have been underestimated or that some or all of the data contain systematic errors. Wherever applicable, we shall quote the larger uncertainty. Cohen and Taylor, J. Phys. Chem. Ref. Data 2, 663 (1973), footnote 8, on the page where section 11 begins

Now read the caption of the table those five reference values come from.

Our final recommended set of constants based on adjustment No. 41 of table 31.2. χ² = 14.50 for 27 − 6 = 21 degrees of freedom; RB = 0.83 the same paper, caption to Table 33.1

The square root of 14.50 over 21 is 0.8309, so the caption's arithmetic checks. And 0.83 is below one, which is the best verdict that diagnostic can return: the inputs scattered slightly less than their own error bars predicted. The table announcing that its data agreed with itself is the table whose values later moved by more than they allowed for.

There is no contradiction there, and that is the point. The Birge ratio asks whether a set of measurements agrees with itself. It cannot see an error they all share. The paper says so itself in the footnote above, where a ratio above one is explained by underestimated errors or systematic error, and where nothing at all is offered for a ratio below one. Consistency is not accuracy, and the decision list on this page is a list of ways to move a consistency statistic. Not one of them reaches the thing that actually went wrong.

Which is the finding The Charge That Crept reached from the other end: the oil-drop measurements of the electron charge agreed beautifully with each other for twenty years, because they shared a wrong value for the viscosity of air.

The same instrument on four more layers

Each of these is a different shape of the same question. In each, one thing moves with the decisions and one thing does not, and the second is the only part worth publishing.

layerwhat the decisions decidewhat survives them
computing

Gold-196, and the rule that will not sit still

The half-life census gives 45 sets of measurements that a real evaluation flagged as contested. Run each set through six combination rules, each crossed with three values of each of the two thresholds the evaluators use, 54 settings in all. The threshold axes are sensitivity probes rather than published rules, and they turn out to be nearly inert: across all 54 settings any one entry yields at most 5 distinct answers, because the threshold rule can only ever hand back one of the estimators the other five already give. Ask how far the recommended value travels, measured in units of the error you would have printed beside it had you never thought about the rule.

median drift
largest drift
further than their own sigma

On gold-196, the case that layer opens on, the three measurements give a Birge ratio of against the the evaluation prints. Across the 54 rules its value runs to days and its uncertainty runs to , a factor of . The Particle Data Group's own scale factor for those same three numbers is , computed from of the 3: its published rule excludes measurements whose error exceeds three root-N times the error of the mean, and the least precise of the three falls outside that ceiling.

A null worth keeping: NUBASE's fixed threshold of 2.5 and CODATA's 1 + √(2/ν) call the same of 45 sets inconsistent and disagree on none. Not every decision moves the answer. This one does not, and saying so is part of the same job as saying when they do.

And a defect in a rule rather than in the data: for of the 45 sets the median-with-scaled-deviation estimator returns an uncertainty of exactly zero, because more than half the measurements share a value and the median absolute deviation is then zero. It is reported here rather than quietly dropped.

Where the plus-or-minus is not there to decide about

Two of the seven layers cannot be run through any of this, and they are the reason the list starts at zero rather than one.

The Room Inside a Number carries intervals that come from rounding rather than from scatter. A printed figure of 0.124 is a claim about the interval from 0.1235 to 0.1245, whose endpoints are exact rationals and whose width the printer chose. There is no sigma on that layer, none is invented here, and every decision above is simply undefined for it.

The Answer That Can't Be Wrong carries integer counts and a category that has to be scored before a rate exists at all. On the same 2,842 bullet comparisons, four conventions that are all in print give false-positive rates of 0.70%, 2.04%, 0.70% and 66.19%. Its uncertainty is a Clopper-Pearson interval on a count, not a symmetric sigma, and running it through a weighted mean would mean inventing a sigma the study never quoted, which is the exact move this page is about.

So the decision before all the others is whether there is a plus-or-minus to decide about. Two of seven answer no, for two unrelated reasons, and they are more useful to this argument than a sixth dataset would have been.

What this page is claiming, and what it is not

It is not claiming that a disagreement is never real. Gold-196's three measurements really do not overlap, under every rule on the grid. Every recommended value of the electron charge from 1929 to 1963 really does sit below the exact value, under every reading of every sign. The historical intervals really did fail to cover, by a factor of at least under all 144 rules. Those are findings, and they are findings because they survived the attempt to move them.

It is claiming that the sentence this is a three-sigma result is not a report about nature until the six decisions are on the table, and that the honest form of any such claim is the range it takes across the decisions you could defend. Where that range crosses a verdict boundary you have learned about your rule. Where it does not, you have learned about the world.

And underneath all six sits the limit none of them reaches. Every decision here moves a statistic that measures whether measurements agree with each other. A systematic error they all share moves the measurements and leaves the statistic alone, which is why a Birge ratio of 0.83 sat in the caption of a table whose values were about to be found wanting. The decisions are worth making carefully. They were never the thing that was going to save you.

Show the check

Everything above is recomputed in your browser from data files committed beside the member layers, and independently in a verifier that does not import the page's engine.

Offline: 142 checks. Run node research/every-plus-or-minus/verify.mjs.

The full check trail, including the two things this page could not settle, is research/every-plus-or-minus/facts.md.