Ground truth · forensic science · signal detection

The Answer That Can't Be Wrong

In 2021 a team at the Ames Laboratory asked 173 firearms examiners to compare bullets they had never seen. On the pairs that did not match, the study reports a false-positive rate of 0.70 percent. The same 2,842 comparisons, in a paper by two statisticians reading the same table, come to 66.19 percent. Neither side miscounted. The gap is one word.

0.70%
as the study reports it
(abstentions counted as non-errors)
the same
2,842
comparisons
66.19%
as its statistician critics report it
(abstentions counted as potential errors)

Both numbers are recomputed on this page from the study's own published counts, which you can see and check below. They differ by a factor of 94.

Three answers, not two

A forensic examiner comparing two bullets is not answering yes or no. The range of conclusions adopted by the Association of Firearm and Tool Mark Examiners in 1992 permits four answers: identification, inconclusive, elimination, and unsuitable. The middle one has three sub-cases, in AFTE's own words:

(a) Some agreement of individual characteristics and all discernible class characteristics, but insufficient for an identification. (b) Agreement of all discernible class characteristics without agreement or disagreement of individual characteristics due to an absence, insufficiency, or lack of reproducibility. (c) Agreement of all discernible class characteristics and disagreement of individual characteristics, but insufficient for an elimination. AFTE Range of Conclusions, adopted April 1992. The labels A, B and C used in the Ames studies and throughout this page are the studies' convention, not AFTE's own.

That third answer is reasonable. An examiner who cannot see enough detail to decide should be able to say so, and forcing a guess would be worse. But it creates a problem that has nothing to do with bullets and everything to do with arithmetic: when you go to compute an error rate, you have a category of answers that are neither right nor wrong, and you must decide what to do with them.

There are four things you can do, and all four appear in the published literature. The framework below is the one set out by Hofmann, Vanderplas and Carriquiry and tabulated against the Ames data by Dorfman and Valliant; the names are theirs.

Instrument 1 · the same study, four ways

The study's raw counts, by true state of the pair and answer given
The false-positive and false-negative rate under each of four published conventions

Every rate in that second table is computed here, in your browser, from the integer counts in the first, by the four formulas above. Where a published source states one of them, the page prints the source's own sentence next to it. Thirteen published figures from five studies and three independent sets of authors are reproduced this way; the offline notebook checks all thirteen, and refuses to build the page if any drifts.

The ordering is not an accident

Look at the second table again and you will see the four numbers never scramble. Writing a for the share of different-source pairs called an identification, b for the share called inconclusive, and c for the share correctly eliminated, the three distinct false-positive rates are

a  ≤  a/(a+c)  ≤  a+b

and that ordering holds always, for every possible study. The proof is three lines. Since a+b+c = 1, the right-hand quantity is 1-c, and the middle one is 1 - c/(a+c). So the right-hand inequality says c/(a+c) ≥ c, which is true because a+c ≤ 1. The left-hand one is the same fact read the other way. The notebook checks it on two hundred thousand randomly generated studies and finds no violation, which is not the proof but would have caught a mistake in it.

So the dispute is not really about which number is right. It is about where in a fixed, guaranteed interval a discipline gets to place its public figure. And the width of that interval is set by one thing: how often the examiners abstained.

What PCAST said, and the trap it names

In 2016 the President's Council of Advisors on Science and Technology reviewed the feature-comparison disciplines and gave an explicit instruction about this arithmetic:

PCAST's worked example is worth running, because it is the whole argument in one line. A method tested a thousand times returns 990 inconclusives, 10 false positives, and not a single correct result. Reported one way it has a 1 percent error rate. Reported the other way, every conclusion it ever reached was wrong.

PCAST's example, live

A method with no demonstrated ability whatsoever, described by the first of those numbers, sounds better than every discipline in the table above.

Why a rate that forgives abstention rewards timidity

PCAST's example is extreme by design. The uncomfortable part is that it is not a special case. It is the endpoint of a smooth process that operates on every real examiner, and the instrument below lets you drive it.

Model an examiner the way psychophysics models any detector. Comparing two items produces a similarity score. When the two really come from different sources the score is drawn from one distribution; when they come from the same source it is drawn from another, shifted to the right by an amount d′ that measures how good the examiner actually is. The examiner then picks two thresholds: identify above the upper one, eliminate below the lower one, and say inconclusive in between.

The width of that middle band is the examiner's willingness to answer. It has nothing to do with their skill. Drag it and watch what happens to the reported rate.

Instrument 2 · the examiner's room

different source same source the inconclusive band

false-positive rate, same examiner, four conventions

Three things are worth noticing while you drag.

The examiner's ability never changed. The two distributions stayed exactly where they were. Only the band moved. Yet three of the four published conventions reported a steadily better number, because three of the four charge nothing for an abstention.

The improvement is unbounded. For any positive ability at all, however poor, widening the band drives the conclusive-only rate to zero. The reason is the ratio of two Gaussian tails: with the band at d′/2 ± w, the false identifications and the correct eliminations both vanish, but the false ones vanish faster, in the ratio e-w d′. The notebook confirms that decay law numerically to within six percent in the tail. An examiner with d′ = 1, which is genuinely poor, can report a one-percent false-positive rate. They need only decline to answer 99.99 percent of the time.

Nothing about this requires anyone to cheat. The Ames I report records that of its 218 examiners, 96 never used the inconclusive category at all, while 45 used it for every single different-source comparison they were given. Those two groups were answering the same questions in the same study, and their personal error rates are not comparable quantities.

None of this is a new observation, and it should not be presented as one. The reductio at the end of the slider was stated plainly by Dror and Scurich in 2020:

If one refuses a priori to count inconclusive decisions as errors, then error rates may be artificially and falsely reduced by making inconclusive decisions. In fact, zero error rates are possible with such an approach: regardless of anything, just reach inconclusive decisions for every comparison and you will have a perfect score! Dror & Scurich, Forensic Science International: Synergy 2 (2020) 333-338

Hofmann, Carriquiry and Vanderplas made the same point independently in the same year, in the paper that named the four options: "an examiner could report inconclusive decisions for every evaluation over the rest of their career and never make an error." What the instrument above adds is only that the effect is continuous rather than a corner case, which is also conceded by the argument's most careful opponents. Biedermann and Kotsoglou, who reject the claim that an inconclusive can be an error at all, nevertheless write that "any less than extreme use of 'inconclusive' also tends to artificially decrease the error rate."

One distinction is worth keeping straight, because it is easy to blur. The reductio above is sharpest against counting abstentions as correct. Against excluding them the arithmetic can run the other way, and at a fixed band it always does: excluding shrinks the denominator, so the ordering theorem guarantees the excluded rate is the higher of the two. Both things are true at once. Excluding abstentions raises the rate compared with forgiving them, and widening the band lowers it under either.

Instrument 3 · how accurate would you like to look?

Fix an examiner's true ability, name the false-positive rate you would like to publish, and the page solves for the inconclusive band that delivers it, then tells you the price in comparisons you must decline to answer.

Two examiners, ranked backwards

The consequence that matters for a courtroom is that these numbers cannot be compared between examiners, or between laboratories, or between disciplines. Here are two examiners. A is markedly better at telling matching bullets from non-matching ones. B is poor but reticent.

Instrument 4 · the inversion

B publishes the better false-positive rate and the better false-negative rate at the same time, on both counts, while being the worse examiner by every measure of actual discrimination. At every matched false-positive rate, A detects more true matches than B does; A's curve lies above B's everywhere. The reported rates have simply stopped measuring the thing their name suggests.

The case against everything above

This is a live dispute among serious people, and the side that says an abstention is not an error has an argument that does not reduce to self-interest. It deserves to be stated in its own words rather than paraphrased into weakness.

The conceptual objection is Biedermann and Kotsoglou's, and it is not about arithmetic at all:

Thus, attempting to label the consequences of 'inconclusive' decisions as correct (accurate) or incorrect (erroneous) is a contradiction in terms when accuracy denotes congruence with ground truth. 'Inconclusives' explicitly make no statement about ground truth. Biedermann & Kotsoglou, Forensic Science International: Synergy 3 (2021) 100147

The practitioners' version, from the Ames II authors themselves, adds the point that the abstention is doing protective work:

Inconclusive decisions are not systematic errors; rather, they are an essential part of the firearms discipline and Inconclusive decisions provide a check against bad Identifications. Monson, Smith & Peters, Journal of Forensic Sciences 68(1) (2023) 86-100

That is a real point and the instruments above support it, not against it: raising the identification threshold does make each identification more probative, and an examiner who abstains rather than guessing is behaving well. A metric that punishes them for it is a bad metric.

Even Dorfman and Valliant, whose recomputation supplies the 66 percent figure this page opened with, do not actually claim that abstentions are errors. Their position is narrower and, in its own way, more damaging:

However, in this paper, we premise that by and large inconclusives in firearms casework reflect the limits of the methods and material, and should not be regarded as errors, that inconclusives in casework are analogous to a medical diagnostic that comes up borderline and fails to indicate whether or not the patient has a suspected condition. Dorfman & Valliant, Forensic Science International: Synergy 5 (2022) 100273

Their argument is about what a study can establish, not about what an examiner did wrong, and they draw the conclusion this page would also draw: not that one number is right, but that no single number is available. "All we can then properly speak of is the potential error rates, which can be assumed to lie somewhere between the minimum and the maximum."

So the honest description of the state of the field is not that one camp is fooling itself. It is that a three-answer procedure is being squeezed into a two-answer statistic, and the squeeze can be performed in four defensible ways that disagree by up to a factor of ninety four. Everyone is computing correctly. The quantity is underdetermined.

The quantity that survives

It would be a thin piece of work to take something apart and leave it in pieces, so here is the part that holds. Two quantities in this picture are immune to the whole dispute, because neither one needs a convention.

The first is the likelihood ratio of the actual report: how much more probable this examiner's stated conclusion is when the pair truly matches than when it does not. It is a property of the report, not of the scoring scheme. Notice, in the panel below, that the value of an identification does not depend at all on where the examiner set their elimination threshold, and that raising the identification threshold makes each identification worth strictly more. Caution is genuinely informative. It just cannot be cashed in as a low error rate.

The second is the whole curve rather than one point on it. An error rate is a single operating point chosen by the examiner's temperament; the curve is the ability underneath. That is what ranks A above B, and it is what a study should report.

Instrument 5 · what does not depend on the convention

Driven by the same two sliders as Instrument 2 above; scroll back to move them.

And there is one more fact, small and rather beautiful, that the model hands over for free. When the inconclusive band sits symmetrically about the point where the two distributions cross, the probability of an inconclusive is exactly the same whether the pair matches or not. Its likelihood ratio is exactly 1. It carries no information at all, in either direction.

That is a real result about the argument this page started with. In the symmetric case, the long dispute over whether an inconclusive should be scored as correct or as an error is a dispute about how to score a report that says nothing. Neither side is right, because the question has no answer: the correct treatment of a genuinely uninformative abstention is to report how many there were and price them at 1.

Slide the band off centre in Instrument 2 and watch the inconclusive likelihood ratio leave 1. That is the case worth worrying about, and it is the case the real data looks like. In the Ames II bullet round, examiners called 20.5 percent of true matches inconclusive and 65.5 percent of true non-matches inconclusive. Those are very different numbers, so in that study "inconclusive" was not a neutral abstention. It was carrying evidence, unlabelled, in a direction nobody had priced.

What this model is not

The examiner in Instruments 2 through 5 is a cartoon: two equal-variance normal distributions and a pair of thresholds. Real examiners do not have a scalar similarity score, real comparison difficulty varies enormously from pair to pair, real inconclusive decisions are shaped by laboratory policy (the Ames II authors report that some participants' laboratories forbid an elimination without knowing the time gap between the collection of the questioned and known specimens), and the three AFTE inconclusive levels are not one undifferentiated band.

So no number produced by the cartoon is a claim about any real examiner. What the cartoon is for is the shape of the effect, and the shape does not depend on the cartoon being right. The ordering theorem, the four conventions, and the replication of thirteen published figures are all pure arithmetic on counts, and they hold whatever generates the counts.

Two further honesty notes. The Ames II technical report ISTR-5220 was withdrawn from the internet by the Ames Laboratory, so the counts used here are taken from the study team's own preprint of the accuracy paper, not from the withdrawn report. And the headline rates in the published journal versions of both Ames studies are beta-binomial maximum-likelihood estimates that account for variation between examiners, not the raw proportions; those are a different and defensible quantity, and this page replicates the raw proportions only, and says so in each case.

What the disciplines were doing meanwhile

The arithmetic above is not an abstract worry. The reason feature-comparison error rates are contested at all is that for most of the twentieth century there were none, and testimony went ahead without them. The National Research Council said so in 2009 in a sentence that has been quoted in courtrooms ever since:

With the exception of nuclear DNA analysis, however, no forensic method has been rigorously shown to have the capacity to consistently, and with a high degree of certainty, demonstrate a connection between evidence and a specific individual or source. National Research Council, Strengthening Forensic Science in the United States: A Path Forward (2009), p. 7

What that absence permitted is on the record. In 2002 FBI scientists rechecked their own microscopic hair comparisons against mitochondrial DNA: of 80 hairs the laboratory had found microscopically indistinguishable, 9 came from different people. In April 2015 the FBI and the Department of Justice, jointly with the Innocence Project and the National Association of Criminal Defense Lawyers, announced the result of reviewing their own hair testimony:

In the 268 cases where examiners provided testimony used to inculpate a defendant at trial, erroneous statements were made in 257 (96 percent) of the cases. Defendants in at least 35 of these cases received the death penalty and errors were identified in 33 (94 percent) of those cases. Nine of these defendants have already been executed and five died of other causes while on death row. FBI national press release, 20 April 2015

Three kinds of statement were counted as erroneous, per the FBI's own review criteria: saying that a hair could be associated with one individual to the exclusion of all others; attaching any statistical weight or probability to a hair association; and citing the examiner's own casework counts as though they were a population frequency. All three are the same underlying mistake, and it is the one this page is about. They are attempts to put a number on a comparison when no study had ever measured what that number was.

The black-box studies are the correction to that. They exist because PCAST and the NRC demanded them, and both Ames studies are serious, expensive, well-designed work by people doing exactly what was asked. That is precisely why the reporting question matters so much now. Having finally measured something, the field has to decide what the measurement says, and at present the same measurement is being quoted in court as 0.70 percent and in the statistical literature as 66 percent.

The thing to ask for

There is a clean answer available, and it is not to pick a convention. Every convention throws away information, which is why they disagree. The full two-by-three table does not: every one of the four rates can be recomputed from it, along with the likelihood ratios, the curve, and anything a future argument might want. Both Ames studies published theirs, which is why this page could recompute them at all, and that is entirely to their credit.

This is not an original prescription, and versions of it are already on the table from people inside the argument. Hofmann, Carriquiry and Vanderplas propose reporting the examiner's error rate and the process's error rate as two separate numbers, since the two questions are genuinely different, and note that a simpler alternative is to report the conditional error rate of each decision: given that the examiner said identification, how often is the pair from different sources? For Ames I that number is 22 in 1,097, or 2.0 percent, and it needs no convention at all. A NIST group reached the neighbouring conclusion from the other direction:

Thus, for non-binary conclusion scales, error rates alone do not provide sufficient information for characterizing method performance (i.e., discriminability and reproducibility). Swofford et al., Forensic Science International: Synergy 8 (2024) 100472

So the request is small and already half-agreed: report the table, report the abstention rate beside every error rate, and never state a forensic error rate without saying which of the four it is. Then the argument becomes one anybody can have with the numbers in front of them, instead of one where two people quote figures ninety-four times apart and both are telling the truth.

Show the check

Every number on this page comes from one file, engine.js, which your browser and the offline notebook both load unchanged. There is no second implementation to drift from.

The notebook, research/the-answer-that-cant-be-wrong/notebook.mjs, runs 103 checks and does not trust the engine anywhere: erfc is verified against values generated by the platform C library through Python, the outcome tables against 400,000-trial Monte Carlo simulation, the ordering theorem against 200,000 random studies, the area under the curve against 300,000 simulated pairs, and the four conventions against exact integer arithmetic. Then it replicates the 13 published rates quoted on this page, drawn from 5 studies and 3 independent sets of authors, from their raw counts. It then does the same for a fourth team, Hofmann, Carriquiry and Vanderplas, whose recount of Ames I differs slightly from the Ames report's own, and shows that the entire gap between 34.76 percent and 34.82 percent is the two blank responses one of them counts as inconclusive and the other does not. Finally it reproduces PCAST's three confidence bounds with an independently implemented Clopper-Pearson interval.

Run it yourself: node research/the-answer-that-cant-be-wrong/notebook.mjs. A second verifier drives this page in a real browser and checks that what you see here agrees with what the notebook computed: node verify-the-answer-that-cant-be-wrong.mjs.

Sources

    Where a source's own text is quoted it is quoted verbatim, including its punctuation. Two small discrepancies between official documents are noted rather than smoothed over: the 2009 NRC report describes the FBI hair result as "9 of them (12.5 percent)" where 9 of 80 is 11.25 percent, and the Ames I technical report says 96 examiners never used the inconclusive category where the 2023 journal version of the same study says 102. Neither affects anything computed here.