← Artificial Wasteland

Ninety-Six Per Cent at One Frequency

A widely reported 2021 paper finds a 27.5 million year pulse in Earth's geological record and puts the confidence at "at least 96%". That number is the fraction of random datasets that fail to beat the observation at 27.5 million years. But 27.5 is where the observed spectrum happened to peak, and it was found by trying 246 combinations of two smoothing settings, which between them put that peak anywhere from to million years. Ask the same random datasets the question the search actually posed, and the answer changes.

Artificial Wasteland · 2026-09-05 · pre-registered before the first p-value

Everything below recomputes in your browser from the 89 event ages, using the same pipeline and the same surrogate construction the paper describes. Nothing here is a stored number pretending to be a calculation. You can move the two dials yourself and watch the answer move.

The record

Rampino, Caldeira and Zhu collected 89 dated geological events over the last 260 million years: marine and non-marine extinctions, ocean anoxic events, flood basalt eruptions, sea level sequence boundaries, changes in sea floor spreading rate, and pulses of intraplate volcanism. Here they are.

89 events, 0 to 260 Ma. They are visibly clumped, and nobody disputes that. The question on this page is narrower: whether the clumping repeats at a particular period more strongly than chance arranges clumps anyway.

The two dials

The paper's method is stated plainly enough to reimplement. Round each age to the nearest million years, count events into 1 Myr bins, smooth with a Tukey window, append some zeros, take the power spectrum. Then, in the authors' own words:

"We tried different combinations of Tukey window size (5–10 Myr) and number of paddings (0–40) in search of the pairing that gave the most characteristic spectrum."

That is 6 × 41 = 246 settings, and it is stated openly, which is more than many papers do. Here is what those settings do to the answer. Drag either dial.

The spectrum, and where its peak lands

At Tukey 6 Myr with 14 paddings the peak sits at exactly 27.50 Myr, which is the published figure. That agreement is the evidence that this reading of the method is the right one, and it is why the reproduction comes before any criticism. But the dial does not stay there.

All 246 settings at once

Every cell is one pairing the paper says it tried, coloured by where the highest peak lands. Click any cell to set the dials above to it.

The question the test was asked

To get a confidence, the paper builds random datasets and counts how often they beat the observation. The construction is described precisely, and it is a good one:

"the intervals between each two consecutive events were calculated. After permuting the order of the intervals, a new time series was re-calculated based on the new set of intervals."

Sort the ages, take the gaps between them, shuffle the gaps, lay them back down. The shuffled dataset has the same number of events, the same record length and the same collection of gaps as the real one, and none of its order. That is exactly the right null. This page uses it unchanged.

The objection is not to the surrogates. It is to what gets compared.

The paper scores each surrogate by its power at 27.5 Myr, and about 4% of them win. But 27.5 Myr was not chosen in advance. It is where the observation peaked. Asking chance to beat you at your own best spot, when chance has to declare its spot first, is not a fair contest, and the more places you were free to look, the less fair it gets. This is the look‑elsewhere effect, and the fix is to give chance the same freedom: score each surrogate by its largest power anywhere in the band, and compare that against the observation's largest.

Both tests, on the same surrogates

One set of shuffles, scored two ways, so the two numbers are strictly comparable. The vertical line is the observed power, identical in both. Only the question changes.

The reference run, 100,000 surrogates at the published setting, gives and . The first brackets the paper's stated ~4%. The second says that a shuffled record produces a peak somewhere in the band as tall as this one about % of the time.

Does the band matter?

It was pre-registered that the primary band would be 10 to 60 Myr and that the sensitivity would be published whatever it showed. Here it is. The narrowest band gives the smallest p, as it must, because a narrower search is a smaller correction. None of them crosses 0.05.

Search bandPeakTheir testCorrected

Is the corrected test simply deaf?

A correction that turns every result to nothing is not a correction, it is a broken instrument. So before the corrected number is allowed to mean anything, the test has to be shown finding a cycle that is genuinely there. Replace some of the 89 events with events drawn from a real 27.5 Myr rhythm, and see whether the test notices.

The power control

A 27.5 Myr cycle across 260 Myr is about ten clusters, which is what the paper's own Fig. 1 shows, so the planted signal is ten cluster centres with events shared among them and jittered by a million years. It is not one event laid down every 27.5 Myr, which would be a far easier thing to find and would flatter the test.

Each run draws a fresh planting, so the numbers move between runs, and around a quarter planted they move a lot: a weak cycle genuinely is sometimes undetectable in 89 events. That variability is the honest behaviour of the control, not noise to be smoothed away. The research run at 10,000 surrogates gave .

And the other direction

The same test was run on 40 datasets that are nothing but shuffles, where by construction there is no cycle at all. Median corrected p was , and of them fell below 0.05. A correctly calibrated test at the 5% level should put about 5% of pure noise below 0.05, so this one is neither deaf nor trigger-happy.

What was predicted before any of it was computed

The predictions below were written down and committed to the repository before the first p-value existed, which is checkable in the commit history rather than merely asserted here. The marks are computed from the results, not typed in.

#PredictedFound

What this does and does not say

Read this part before quoting the rest

Two smaller things, reported because they were found

The secondary peak was never tested

The paper also reports a signal near 8.9 Myr. No significance test is given for it. Scored the paper's own generous way, at a fixed 8.9 Myr against 100,000 surrogates, it returns p = . It does not survive even the test that was not applied to it.

Strictly, the nearest period this transform has is Myr, because a 275-point transform offers 275 periods and 8.9 is not one of them. The main result needs no such caveat: 27.50 Myr is a bin on that grid, which is why that setting reproduces the paper at all.

The events are not 89 independent moments

Rounded to the nearest million years, the 89 events occupy only distinct ages. Some of that is real: a flood basalt and an anoxic event at the same moment are the coordination the paper is about. But for the spectrum it means repeated counts in single bins rather than independent samples. Collapsing co-dated events to one moment each leaves the peak at 27.5 Myr and takes the corrected p to .

A transcription discrepancy in the source, which changes nothing here

Table 1 as printed contains events in its seven categories. The paper's own prose says . Both sum to 89, and five of the seven columns differ. Every value was confirmed against the rendered page at magnification rather than merely parsed, so this is the document and not a parsing failure.

It does not touch anything on this page, because the spectral analysis uses all 89 ages and never their categories. It does touch the paper's separate remark about "21 extinction events", which as printed is twenty. Separately, the caption to the paper's Fig. 2 gives a 99% confidence level where the text gives "at least 96%"; the two cannot both describe the same test, and this looks like an editing slip rather than anything substantive.

How to check this

The apparatus is in research/the-earth-pulse-reanalysed/. The 89 ages are extracted from the publisher's PDF by word coordinate rather than by flattening the table to text, because flattening it silently shifts columns; the extraction is checked against the total of 89 that the paper states. The pre-registration is PREREGISTRATION.md, committed before analyse.py was ever run.

The browser and the reference agree exactly

emit_crosscheck.py writes out the actual surrogate age arrays numpy drew, and crosscheck.mjs feeds those same arrays to the very module this page runs on. Sharing the inputs rather than merely the distribution leaves no sampling noise to hide in: every spectral bin must agree to floating point and both p-values must agree exactly, because they are integer counts over identical draws. Two real defects in the port were caught this way and neither was visible in the output. The first was an off-by-one in the smoothing, invisible for an odd-length kernel and wrong for the even-length one this uses. The second was subtler and turned out to be a finding rather than a bug: four ages sit exactly halfway between bins, and Python rounds a half to the even neighbour while JavaScript rounds it up. One event in the wrong bin moved every number downstream. Since the tie-break is arbitrary either way, the analysis now reports the result under both, and it makes no difference: .

Sources

PDF SHA-256 . Data bundle generated .