One dataset, every defensible verdict

Conformity
Is a Choice

A paper rejects Benford conformity for 19,509 US city populations. Another published standard calls the same rows close conformity. Both calculations are right. The verdict is coming from somewhere else, and this page goes and gets it.

Who decided this dataset fails?

This page assumes you know the digit law. The First Digit Is a One explains it. Most Numbers Begin With One currently says a dataset “conforms at 5% if chi-squared is below 15.51.” Here we go after the sentence itself: where the competing cutoff came from, what its author said it does, and whether that is true.

Loading two embedded Census vintages and checking the ten-number anchor...

The same rows, opposite labels

Kossovsky’s 2021 paper prints chi-squared = 17.4 > 15.5 and rejects. Reading the same first digits through Nigrini’s fixed MAD band gives MAD = ..., inside the close-conformity range. A third published statistic, Cerqueti and Lupi’s excess MAD, puts the same digits ... null standard deviations from Benford, which is a rejection again.

Pearson chi-squared, alpha 0.05 REJECT

The paper prints .... Exact logarithmic probabilities recompute ..., with exact tail probability ... on 8 degrees of freedom.

Nigrini first-digit MAD band CLOSE

The same ... observations give MAD = ..., under the 0.006 close-conformity line.

paper prints17.4chi-squared
4-decimal probabilities...prints 17.4
exact logarithms17.5236prints 17.5
exact SSD...prints 1.3

The gap between those two recomputations is the precision at which you write down the Benford probabilities before you subtract them. Rounding the null vector to three, four or five decimals moves the statistic across a range of ... without moving the verdict.

null vector precisionchi-squaredverdict at 0.05

The paper’s printed 17.4 is matched only by the four-decimal row, and that convention does not reproduce four of the paper’s other five printed first-digit values. We do not know which precision was used. That is the first undocumented choice on the page and it is one of the smallest.

Six datasets, six disagreements

It is not one lucky table. These are all six datasets the paper analyses, with their first digits re-extracted from the original workbooks. Chi-squared rejects .... Nigrini’s band calls ... close or acceptable. Kossovsky’s own band calls ... approximately perfect or acceptably close. Nothing about the digits changed between columns.

datasetnchi-squaredpat 0.05MADNigrini 2012SSDKossovsky 2021

A second published anchor, on the same six files: Cerqueti and Lupi convert MAD into a sample-size-aware z. Five of their six printed values reproduce exactly here from a covariance derivation written for this page.

datasetprinted zrecomputed zresidual

The two residuals are the page’s own subject showing up inside its anchor.

Oklahoma. The paper describes its rows as positive payments under one million dollars and prints n = .... That count includes ... payments of exactly zero, which have no first digit at all. Strictly positive gives ..., which is exactly the n the other paper prints. One file, one stated rule, two sample sizes, over whether zero counts as positive.

Bone marrow. The gap closes when ... values at or below 1e-06 are dropped. They sit on a 1e-07 lattice, a detection floor no workbook mentions. At n = ... the recomputed z is ... against a printed 5.034. Somebody applied a floor and did not write it down.

Where the cutoffs came from

The 0.006 line that just acquitted this data is the most-cited number in Benford-based forensic accounting. Here is its entire published derivation, in its author’s words.

“The MAD results in Figure 7.5 were used to create a set of critical values to update the table in Drake and Nigrini (2000). These new MAD critical values are shown in Table 7.1.”Nigrini, Benford’s Law: Applications for Forensic Accounting, Auditing, and Fraud Detection, Wiley 2012, chapter 7
“Drake and Nigrini (2000) offer some guidelines based on personal [experience]”same chapter, describing the source it updates

That is the chain. A figure of MADs computed on datasets the author considered conforming, replacing an earlier table offered on personal experience. There is no distributional argument anywhere in it, and no statement of the sample sizes the calibrating datasets had. We have not seen Figure 7.5 or the 2000 table, so we do not print either, and the current numbers should be attributed to Nigrini 2012 alone.

Here is the table itself, verified word for word from the book’s own scan.

digit testclose conformityacceptablemarginally acceptablenonconformity

Rows marked * rest on a single optical-character witness from that scan. The first-digit and first-two-digit rows have an independent numerical confirmation: inverting a third paper’s published excess-MAD offsets recovers 0.005996 and 0.001197.

The competing standard is franker about itself. Kossovsky prints his own SSD ladder and then this:

“These rough guidelines however were subjectively arrived at, albeit with outmost effort to make them as reasonable as humanly possible.”Kossovsky, “On the Mistaken Use of the Chi-Square Test in Benford’s Law”, Stats 4(2), 2021, page 445
“Since SSD does not incorporate the term N in its expression, one gets the same measure and conclusion regardless of the number of observations. But this comes with a certain price, as there is no associated statistical theory to guide us. There is also no hope that future statistical studies would somehow yield threshold points, significant values, or confidence intervals by applying SSD, since those highly beneficial results would certainly require involving N somewhere in the relevant expression, while N is nowhere to be found in the definition of SSD.”same paper, section 12
SSD testapproximately perfectacceptably closemarginally Benfordnon-Benford

He makes the symmetric point too, and it is fair: the 5% and 1% probabilities attached to chi-squared are also conventions nobody derived. The difference is what happens to each convention when the sample size changes, which is the next section.

The reason given for the test is false

Both fixed-band standards exist to escape a real problem: chi-squared convicts large samples for deviations too small to care about. Here is the premise offered for the escape.

“What is needed is a test that ignores the number of records. The mean absolute deviation (MAD) test is such a test.”Nigrini 2012, page 158

MAD does not ignore the number of records. Take a sample of size n drawn from the exact Benford distribution, so that nothing whatever is wrong with it, and ask what each statistic is expected to return.

digit windowexpected chi-squaredexpected MADexpected SSD

Chi-squared alone has no n in it. Its expectation is the degrees of freedom whatever the sample size, which is why it can carry a fixed critical value. The expected MAD of perfect data falls as ... over the square root of n, and the expected SSD falls as ... over n. Both fixed ladders are therefore ladders in disguise for sample size, and both were calibrated at whatever sizes their calibrating datasets happened to have.

Which turns the bands into arithmetic about n. Below these sample sizes, perfectly Benford data is expected to be labelled by the band shown, no matter what the data does.

digit windownonconformity cutoffexpected nonconformity below nclose cutoffclose conformity unreachable below n

Read the first-digit row plainly. A perfectly Benford dataset of fewer than ... records is expected to be declared nonconforming, and no first-digit dataset under ... records can expect to earn close conformity however Benford it is. On the first-two-digits test, the one recommended for invoice data, nonconformity is the expected verdict for perfect data below ..., and at n = 1,000 it happens in ... of samples.

Those are analytic statements. Below they are re-run as exact multinomial draws, which is a check that can disagree with the algebra and does not.

nreplicatesmean chi-squaredchi-squared rejectsmean MADNigrini says nonconformityNigrini says closeKossovsky says non-Benford

Running the exact multinomial null simulations...

The chi-squared column is the control. Its rejection rate sits between ... and ... across every sample size in the table, which is what a correctly sized 5% test looks like once Monte Carlo error is allowed for. The much-repeated “excess power problem” is not a problem with the test’s size at all. It is a statement about behaviour when the data really is not Benford, and it is a legitimate complaint. The fix on offer replaces a test whose error rate is known with a ladder whose meaning changes with every sample size, and says in print that it does not.

Turn one choice

Back to the one dataset. Seven statistics, three digit windows, four significance levels, twelve year columns, three small-place floors, and two official data vintages. The fully crossed grid contains ... specifications, every one of them an alpha-calibrated test. Select any cell. The browser recomputes its statistic and tail probability from the shipped integers, and prints the two fixed-band standards on the same slice beside it.

Computing selected specification...

...... specifications reject
...... specifications conform

The published rejection sits at the ... percentile of the calibrated verdict-margin curve. It is real, but it is the minority conclusion on its own data. Its Monte-Carlo-calibrated tail probability is ... against the exact chi-squared value quoted above, a difference of resolution and not of verdict.

Preparing the full curve...

The two fixed-band standards cannot be plotted on that axis, because they are distances with no tail probability attached, and pretending otherwise would invent a common scale that nobody in the literature has. So they get counted instead. Strip the decision rule out of the grid and ... distinct slices of this dataset remain: window by year by floor by vintage. Here is the share of those slices each published standard calls nonconforming.

Share of slices called nonconforming, by standard

Chi-squared at 0.05 calls ... of the slices nonconforming. Nigrini’s nonconformity line calls .... Kossovsky’s non-Benford line calls .... The published analysis of this dataset used the first of those. That is a fair thing to notice and not an accusation: a paper complaining that chi-squared rejects too much was written by an author whose own specification sits in the most rejection-prone rule available.

A caution that applies to every share on this page, including the two above: it depends on how many variants of each kind we chose to enumerate. Add four more year columns and the year dimension gains weight for no reason connected to the world. Shares over a grid are a description of the grid.

The argument is about the wrong knob

Within the calibrated grid, researchers argue over the test statistic. Marginally, that choice moves rejection by about twelve percentage points. The digit window moves it by about fifty-one, and dropping towns below 100 moves it by almost fifty.

Exact factorial decomposition of the binary verdict

Largest main effect: ..., ... of total verdict variance. Interactions carry .... The complete balanced grid makes these shares exact.

There are two honest summaries. Digit window has the largest top-to-bottom marginal swing. The small-place floor carries the largest sum-of-squares main effect because its three levels separate more persistently across the full factorial grid. Test statistic is near the bottom under either description.

Now change no proportions at all

For fixed digit proportions, Pearson’s chi-squared is exactly linear in sample size. Move the slider. This is a thought experiment that scales the observed nine proportions without changing their shape.

scaled chi-squared...
towns removed...
share removed...

Computing...

At 17,264 observations, 2,245 fewer than the real table, the unchanged proportions cross below 15.507. That is 11.5% of the rows. But “delete 11.5% at random and it flips” would be too strong: random deletion changes the proportions too.

Here is the separate empirical check. For each point the browser takes 400 seeded samples without replacement from the actual first digits and reruns chi-squared.

Running seeded random subsamples...

At n = 17,265, random subsampling passes in ... of these draws. The deterministic flip belongs to the fixed-proportion thought experiment, not to every random deletion. Put the two directions together and the shape of the field appears: chi-squared tends to convict large samples, and a fixed band tends to acquit them. An analyst who chooses the unit of record, invoice line or invoice, precinct or county, month or year, has chosen the verdict before opening the file.

Break it on purpose

Every control above was chosen to stay out of the regions where the arithmetic stops being about the data. This one is built to go there, and to refuse. Pick a value range for the city populations and watch what the page will and will not say.

rows retained...
decades spanned...... digits impossible
chi-squared...computed regardless
MAD...computed regardless

Computing...

Two gates are running there. The first refuses when the retained values span less than one decade, because then the leading-digit distribution is a property of the filter: restrict to 200 through 600 and four leading digits become arithmetically impossible, and every test in the literature returns overwhelming nonconformity about the slider. A minimum and maximum amount is the single most natural control to put on a page like this, and both are documented standard practice in auditing.

The second is the gate the whole page has been building toward, and it is the one check here that can go red on its own numbers: is the expected MAD of perfect Benford data at this sample size already outside the band we want to claim? Put a floor at 300,000 people and 61 cities remain. At n = 61 the expected MAD of perfect Benford data is ..., more than double the nonconformity cutoff, and an exact multinomial simulation returns nonconformity for genuinely Benford data ... of the time. There is no data you could put in that box that would pass. So the page declines to grade it.

The same gate runs over the whole 6,048-cell grid above. Its smallest reachable sample is ... rows, where perfect Benford data expects MAD ..., comfortably inside close conformity: .... That sentence is computed, not asserted, and it would fail if any floor in the menu were raised.

The third trap is subtler and is a plausible implementation slip rather than a control. The sample-size-aware statistic works by subtracting what MAD would be under the Benford null. Subtract instead what MAD would be under a resample of the observed digits, which sounds like the same sentence, and the excess is zero by construction.

datasetchi-squared pexcess MAD z, null-centredexcess MAD z, self-centred

The star distances are rejected by chi-squared at a tail probability with 110 zeros after the decimal point. Centred on themselves they score .... Nothing about that number is wrong: the resampling distribution is doing exactly what it was asked to do, and what it was asked was meaningless.

What did not reproduce

The paper’s second-digit line for this dataset is 13.1. Four defensible conventions are available, crossing whether one-digit values are dropped or padded with a zero against whether the null vector is exact or printed to three decimals, and they span ... to ....

conventionnchi-squaredverdict at 0.05

One of the four rounds to the printed value. That is not enough to call the convention identified, because the same combination does not reproduce the paper’s printed second-digit numbers on its other datasets. What the table does establish is that an undocumented convention moves this statistic by about nine percent, and that all four conventions conform at the 9-degree-of-freedom critical value of ... while all four first-digit conventions reject. Same file, same paper, opposite direction, depending only on which digit you look at.

The check, recomputed in front of you

What exactly is in the shipped data?

Two public-domain CSV files, each 19,510 rows by 12 population columns. One is extracted from the Census table preserved in Kossovsky’s workbook. The other is extracted from the current Census Vintage 2009 “All States” file, retaining only SUMLEV 162 incorporated places. Names, state codes, place codes, and non-population columns were removed because no test on this page uses them. Missing historical counts marked X in the source are stored as -1 and excluded by the same positive-value rule as zero.

The six-dataset table ships as nine integers per dataset, the first-digit counts, extracted from the six original workbooks by a script in the research directory. The workbooks themselves are large and are not redistributed; their URLs and hashes are recorded there, and the counts are checked in the verifier against two independent published papers.

Methods and sources

The seven statistics are Pearson chi-squared, Kossovsky SSD, Nigrini MAD, Cho and Gaines Euclidean d*, Kolmogorov-Smirnov D, Kuiper V, and discrete Cramer-von Mises W-squared. P-values are upper-tail proportions under a seeded asymptotic multinomial null. The score plotted above is log10(alpha divided by p), so zero is the verdict boundary and positive values reject. The null simulations are exact multinomial draws rather than the Gaussian approximation, because the behaviour being tested is what happens at small n and a Gaussian null would assume the answer.

Anchor: Alex Ely Kossovsky, “On the Mistaken Use of the Chi-Square Test in Benford’s Law”, Stats 4(2), 2021, with the six datasets from the Williams College Benford resources page. Second anchor: Roberto Cerqueti and Claudio Lupi, “Severe testing of Benford’s law”, TEST 32:677-694, 2023, Table 1. The conformity bands are Nigrini, Benford’s Law: Applications for Forensic Accounting, Auditing, and Fraud Detection, Wiley 2012, chapter 7, Table 7.1, and Kossovsky 2021 Figure 25. The seven-statistic comparison follows Cano-Rodriguez, “How Much Is Too Much?”, 2025.

Not established here, and therefore not claimed: the numbers in the 2000 table Nigrini says he replaced; the contents of Figure 7.5, which is the whole stated basis for Table 7.1; and whether the second edition of the book revises any of it. Kossovsky’s printed second-digit figures for the Oklahoma file could not be reproduced by any subset we tried, differing by roughly a factor of eight, so this page cites none of them.