Ground truth · 30 August 2026

The Apparatus Is the Finding

Three real datasets, each with a finding worth a headline, and one question put to each: were the two halves of this comparison produced by different apparatus, so that a good part of what looks like a fact about the world is a fact about the equipment? Twice the answer is yes, and the finding moves or fades once the apparatus is held fixed. The third time it is no: those numbers reproduce to the unit and the explanation this page had for them fails a control this page set for itself, which is the more useful of the two outcomes and is why the panel is still here.

The general shape is old and has a dull name: measurement confounding. Nothing here is a discovery about statistics. What is here is three instances with their numbers re-derived from primary sources tonight, a bench at the end that puts the mechanism under your finger, and one rule that comes out of all of it and that no amount of care with covariates will save you from:

A DIFFERENTIAL CHECK SEES ONLY THE PART OF A FAULT THAT IS NOT COMMON TO BOTH OF ITS SIDES.

That is not the same as ordinary confounding, and the difference is the reason for this page. An ordinary confounder is a column: year, size, region. You can condition on it. An apparatus confound is not in the data at all. It is a property of the machine that turned the world into the data, it is usually collinear with the thing you are comparing, and when it is applied to both arms of a comparison it leaves almost nothing behind for any check that compares the arms to each other. The bench at the end of this page is where the word almost gets measured.


One. The letters that looked more censored

openFDA transparency/crl bulk export · 458 letters · export dated 2026-08-26

When the United States Food and Drug Administration declines to approve a drug it sends the sponsor a Complete Response Letter saying why. In 2025 the agency began publishing them directly, and openFDA now ships every one it has released as a single structured file with the full text of each letter in a column. Some of those letters had been visible for years inside Drugs@FDA approval packages and some had not, which is a fact worth holding on to. This page holds 458 of them: 309 whose application the file marks Approved and 149 marked Unapproved.

Redactions in these letters are marked in the text with the FOIA exemption that justifies them, and the one that matters is exemption 4, trade secrets, written (b)(4). Count those marks, divide by the length of the letter, and a finding falls straight out.

Panel one: redaction density, marks per 1,000 characters

Count a redaction as
Compare

Loading.

On the first setting the answer is stark and it is the answer anyone would have published: the letters for applications the file does not mark as approved carry 0.229 marks per 1,000 characters against 0.076 for the rest, a factor of 2.99. The letters for the drugs that did not get through are the most censored ones. That is the sentence this comparison puts in a headline, and the rest of this section is about where it comes from.

Which batch each letter came out of

The export carries the source PDF's filename, and two families appear. Files named CRL_NDA219398_20260128.pdf come from the 2025 transparency release. Files named 208352_2020_Orig1s000OtherActionLtrs.pdf were lifted out of a Drugs@FDA approval package. There are 162 of the first and 296 of the second, and here is how they line up against the grouping variable:

Letters by approval status and by the machine that produced the text, recounted in your browser.
Approval statusTransparency releaseApproval packageRow total
Approved14295309
Unapproved1481149
Column total162296458

443 of the 458 letters, or 96.7 percent, sit on the diagonal. The grouping variable and the batch are very nearly the same variable. So the comparison in the first panel is not approved against unapproved. It is, 443 times out of 458, one release batch against the other.

Both batches were read off the page by a machine

The obvious story, and the one this page was drafted with, is that the transparency release shipped as PDFs with a text layer while the approval-package letters are photographs of paper. That story is wrong and the file says so. Every letter here opens with the FDA letterhead, and in front of the letterhead sits whatever the reader made of the agency's seal. Count the characters emitted before the first letterhead phrase and 98.8 percent of the transparency-release letters carry some, against 99.0 percent of the approval-package ones. One of them opens i Lk YN U.S. FOOD & DRUG ADMINISTRATION, and it is a transparency-release letter. A second signature agrees: one and two character junk tokens run at 5.63 per 1,000 characters in the release batch against 4.95 in the approval-package batch, a gap of 14 percent, in a pair of arms whose marker densities differ by 5.19 times.

So both arms are the output of something that read a page image, and the difference between them is not whether a machine read the letter. It is how often the redaction marker survived being read. That is still an apparatus confound and it is still 96.7 percent collinear with approval status, but the apparatus is the text pipeline rather than one particular scanner, and this page cannot tell you which step inside it differs.

What the pipeline does to a redaction

A redaction on paper is a black rectangle. Whatever read these pages did not read black rectangles as (b)(4). It read them as whatever glyph they most resemble. Here is one letter, CRL_BLA761400_20240820.pdf, which the marker counter scores at 0 redactions. Read the file name: this is a transparency-release letter, one of the 162 in the arm that was supposed to be the clean one.

During a recent inspection of O@ (FEI: ©) listed in this application, FDA conveyed deficiencies to the representative of the facility.
Verbatim from the text field of that record. The facility name and its FDA establishment identifier have both been removed. The counter that scored this letter at zero was looking for the wrong thing.

So count the other thing as well. A second instrument, written against that one letter and then run unchanged over all 458: a standalone © or ®, or a short run of the letters o, O and @ standing alone, with no word character on either side. It shares no code and no characters with the first instrument. Switch the panel above to the second counting rule and watch what happens to the pipeline comparison.

Two instruments, per 1,000 characters. The per-letter counts were made offline from the letters and embedded in this page; every rate here is divided out in your browser.
Measured withTransparency releaseApproval packageRatio
(b)(4) marks0.2660.0515.19 ×
pipeline glyphs0.0720.2823.95 × the other way
either one0.3370.3331.011 ×

The (b)(4) counter says the transparency release is 5.19 times more redacted than the approval package. The glyph counter says the approval package is 3.95 times more redacted than the release. Add the two instruments together and the two pipelines redact at 0.337 and 0.333 marks per 1,000 characters, a ratio of 1.011. They are the same rate. A permutation test on the batch labels puts that difference at p = 0.945, which is about as close to no difference as 999 relabellings can report. The entire five-fold gap was in how often the marker survived being read, not in how much was taken out.

Run the original comparison with both instruments and the headline does not turn over. It goes away. Approved letters carry 0.354 per 1,000 characters and unapproved letters 0.283, a ratio of 0.80, and the same permutation test on the approval labels returns p = 0.163. That is not an inversion and this page will not sell it as one: drop the twenty letters carrying the densest redaction and the ratio is 1.08. What the file supports is the disappearance of a gap that stood at p = 0.001 on the same test before the second instrument was added.

There is a third cut, and it is here only because leaving it out would be a choice made after seeing it. Inside 2024, the one year both groups occupy, the marker counter alone gives 0.290 for approved against 0.241 for unapproved. But the approved arm there is 14 letters, twelve out of the approval package and two out of the transparency release, and 81 percent of its marks come from those two. Take them out and 2024 reads 0.067 against 0.241, pointing the original way. It is not a second, independent control. It is the batch control again, wearing a year.

PANEL ONE: hold the release batch fixed and the gap disappears, p = 0.945. It was the text pipeline, not the censor.


Two. Nine hundred and seventy-nine structures in two months

Paolo et al., Nature 625 (2024) · deposited global time series, 2,498 bytes · CC BY 4.0

A 2024 Nature paper mapped every fixed structure in the coastal ocean from Sentinel-1 synthetic aperture radar, month by month from January 2017 to January 2022, and deposited the global time series as a 2.5 kilobyte CSV. That file is embedded in this page in full. The paper's own sentence about it, which has been widely repeated:

The number of offshore oil structures has increased by about 16% over the past half a decade (Fig. 4c), with a decrease in the USA of several hundred structures offset by increases elsewhere. Paolo et al., Satellite mapping reveals extensive industrial activity at sea, Nature 625, 2024.

The series opens at 7,164 high-confidence oil structures in January 2017 and closes at 8,389 in December 2021, which is the paper's about 16 percent. Here is the whole file. Drag the anchor to choose which month the five-year growth is measured from.

Panel two: growth in offshore oil structures to December 2021

Loading.

The chart draws the oil and wind columns of the embedded file across all 61 months, with the first two months shaded and the anchor month marked. Every figure quoted below is read from the same rows.

Moving the anchor by two months takes the headline from +17.1 percent to +3.0 percent. Everything the statistic is made of happens before March 2017. The first two months of the record add 979 oil structures, which is 79.9 percent of the entire five-year gain of 1,225. The single largest month is the first, +716, against an interior median absolute monthly change of 25.

Nine hundred and seventy-nine offshore oil structures are not built in two months, and the file says so without needing an outside source: no other two-month window in the record comes near it, put the same fifty-nine monthly changes in a random order and a first-two-month share this large comes up 0.002 of the time in 999 tries, and between December 2017 and December 2021, with the ramp finished, the same detector records 8,464 and 8,389: flat, very slightly down.

The anchor is in the paper's own figure caption

The two arms of this comparison are 2017 and 2021, and they were measured by different amounts of satellite. The paper says so:

Spatial coverage also varies over time and is improved with the addition of S1B in 2016 and the acquisition of more images in later years. Paolo et al. 2024, Methods.
The area of the ocean imaged every day by the Sentinel-1 GRD product (using a 12-day rolling average) depended on whether one satellite was imaging the ocean (S1A, October 2014 to present) or two (S1A and S1B, September 2016 to December 2017). S1B stopped operating on 23 December 2021. Paolo et al. 2024, Extended Data Fig. 1 caption.

That last sentence is a published, dated, physical fact about the instrument, and it makes a prediction this page did not choose after seeing the answer: the record should fall off a cliff at the end, because one of the two satellites stopped. The final row of the file is dated 2022-01-01 and the standing total drops by 623 structures in that one month, against an interior median absolute monthly change of 130. Structures cannot be removed from the sea at that rate any more than they can be installed at it. Both ends of this series are the apparatus.

The third symptom is the unclassified detections. If the instrument's reach were constant, the other class would be roughly constant too. It runs 1,276 to 2,651, more than doubling across the window, which is the same shape as the wind turbines and is not a construction programme.

And the crossing has three dates

The paper's other quoted result is that offshore wind turbines overtook offshore oil structures. The file supports three defensible readings of the phrase "oil structure", and they give three answers:

First month in which the wind count exceeds the oil count, under three definitions, computed live.
Definition of an oil structureWind first exceeds itWindOil
high-confidence oil only2020-10-018,6428,552
oil plus probable plus possible2021-06-019,8299,715
oil_other, the paper’s upper bound2021-12-0111,19211,040

The three answers are 2020-10-01, 2021-06-01, 2021-12-01: fourteen months of spread inside one 2.5 kilobyte file. The paper wrote "probably surpassing the number of oil structures by the end of 2020", and the hedge was carrying weight.

PANEL TWO: 80 percent of five years of growth arrives in the first two months, and the paper's own figure caption names the second satellite whose coverage was still filling in, and the day it was switched off.


Three. The panel that did not survive its own control

GBIF Occurrence API · year facets fetched 2026-08-30

A holotype is the single physical specimen that carries a species name. The Global Biodiversity Information Facility will tell you, in one query, how many plant holotypes were collected in each year of the record. The shape of that answer is arresting. The 1930s hold 19,285 plant holotypes. The 2010s hold 5,100. A fall of 3.78 times.

The reading this page was built to make is that this is not botany. It is the shape of a funding programme: the Mellon Global Plants Initiative paid to photograph historical type specimens in large northern herbaria, so the decades it targeted are visible and the recent decades, whose types sit undigitised in small tropical collections, are not. Two arms, two apparatus, finding explained.

Before believing that, normalise. Divide holotypes by all preserved plant specimens collected in the same decade, which removes any effect of how much collecting happened. Then run the control that discriminates: do the same thing for animals. No botanical digitisation programme touched the animal collections. If the collapse is the Global Plants Initiative, the animal curve should not have it.

Panel three: type specimens per million preserved specimens, by decade

Show

The chart draws the plant and animal curves for whichever of the three views is selected, on one shared scale, across the twelve decades in the table below. Every figure quoted in the text is read from the same table.

Loading.

Holotypes by collection decade in the two kingdoms, recomputed from the pinned GBIF year facets.
DecadePlantsAnimals
1900s13,47814,397
1910s13,14215,530
1920s18,31020,293
1930s19,28525,800
1940s13,39915,451
1950s12,07025,027
1960s14,08836,812
1970s15,44431,530
1980s15,17730,016
1990s11,39424,279
2000s8,96722,529
2010s5,10013,593

The animal curve has it. Normalised, plant types fall from 3,052 per million specimens in the 1930s to 640 in the 2010s. Animal types fall from 5,855 to 746 over exactly the same decades. Indexed to the 1930s the two land at 0.210 and 0.127, and across all twelve decades the correlation between the two logged curves is 0.943. Two curves that both fall the whole way will correlate whatever the reason, so that last number is the weakest one here; the two doing the work are the indices.

A funding programme that never touched zoology cannot explain a curve that zoology has too. Two explanations that need no apparatus at all are sitting there instead, and both are documented: a species described today may have been collected decades ago, so recent collections have not yet had time to become types, and a flora that has been worked over for two centuries yields fewer new species per sheet than a flora that has not. Either one predicts this shape. This page's thesis does not get to claim it.

PANEL THREE: the numbers reproduce exactly and the explanation does not. Reported as a failure, because a panel that failed is worth more than a panel that was asserted.

And the control I just ran has the fault this page is about

Read the last three paragraphs again. The control was a comparison: plants against animals. Anything GBIF's own ingestion did to both kingdoms, which datasets were mobilised in which year and which decades of specimens they happened to contain, is applied to both arms of that control and is therefore invisible to it. What the control rules out is a plant-specific apparatus. It cannot rule out a GBIF-wide one, and no differential I can build out of GBIF can, because everything in GBIF came through GBIF.


The bench

The sophisticated objection to everything above is short and correct as far as it goes:

This is confounding. Every analyst knows to control for confounders. Stratify, add the covariate, run the balance checks, and you are fine.

The model behind that sentence is that a confounder is a column you can condition on, and that a check battery which comes back clean means the comparison is sound. Both halves of that are testable, here, on the real letters, and the second half is false in a specific way that no amount of conditioning repairs.

The bench below takes the 162 letters that came through the transparency release, the ones with an intact text layer, and splits them into two arms by a deterministic rule: sort by length, then alternate. The two arms are drawn from one population, so the true difference between them is whatever chance put there. Then it applies a fault, and the fault is not invented: it is the one measured in panel one. Passing a letter through the pipeline that eats markers turns some fraction of its (b)(4) marks into glyphs the counter does not see, and comparing the two batches gives that fraction as 0.805.

Apply it to one arm. Then apply exactly the same fault to both.

The fault bench: 162 real letters, one measured fault, seven checks

Apply the extraction fault to

Loading.

With the fault on arm A only, the comparison is visibly broken and the battery says so: the arms differ by 157 percent, the permutation test on the arm labels returns p = 0.001, and the split-half check and the absolute anchor both fire. Four of the seven light up: the size of the gap, the permutation test, the split-half check and the anchor. Any analyst would stop.

Now put the same fault on both arms. The measured redaction density in each arm falls by a factor of 6.4. Every number the analysis produces is wrong by that factor. And the battery goes quiet: the arms now differ by 17 percent, which is less than the 25 percent they differed by before anything was done to them, and the permutation test returns p = 0.624 against p = 0.302 on the untouched data. Three of the six differential checks can move under a fault of this shape, and all three move the wrong way: they are passing more comfortably than on clean data, because the fault took variance out of both arms at once. The other three cannot move at all, and their rows say so: arm size, covariate balance and the placebo outcome are identical in all three settings by construction. Six checks, three of them asleep, none of them alarmed.

That is the discriminating prediction, and it is the reason this bench is not a demonstration. Under the model in the objection, a check battery detects faults, so the both-arms row of the battery should look like the one-arm row. Under the model this page is arguing, a differential check detects asymmetry, not error, so the both-arms row should be indistinguishable from the clean row. Flip between the three settings and read the battery. At the fault this page measured, if any of the six differential checks fired on the both-arms setting, this page would be wrong.

Where that stops being true, which is the rest of the rule

Now drag the fault slider, and some of them do fire. That is not a hole in the argument. It is the other half of it, and it is why the rule at the top of this page says the part of a fault that is not common to both sides and not any fault applied to both sides. The fault is applied to each letter as a whole number of marks, round(p x marks). Near the measured value that behaves like a scaling and lands on the two arms alike. Wind it up and it behaves like a threshold: every letter holding one or two marks loses all of them, the survivors pile into a handful of letters, and two arms with different mark distributions stop being faulted identically. What is left over is asymmetry, which is exactly what a differential check is built to see. Of the 201 settings the slider can reach, 98 put at least one differential flag on the both-arms setting. The quiet band around the measured fault runs from 0.750 to 0.830. The check panel below sweeps every setting and prints that band, so it is asserted rather than promised.

One of those flags is worth naming. The split-half check reads 1.95 against a threshold of 2 on untouched data. It sits two and a half percent from firing before anything has been done to either arm, so it will fire on rounding noise, and a reader who watches it go red will be told a fault was caught when none was. That is this page's own disease, one level up, in this page's own instrument.

The seventh check

One row does fire, and it is the only one that is not a comparison between the arms. It puts the number each arm produced next to a level obtained from a different instrument: the mark-or-glyph counter from panel one, which is unaffected by the fault because the glyphs are exactly what the fault produces. On clean data the marker counter reads within a factor of 1.3 of it. Under the both-arms fault it reads a factor of 9.0 away, and says so.

So the remedy is not more controls, and one qualification is owed to the objection before that sentence is allowed to stand: a control did do the work in panel one. Switching that panel to compare the two batches is conditioning on the machine, and it exposed the confound exactly as the objection says it should. Conditioning works whenever the apparatus is a column you can see. The trouble is that when it is not a column, conditioning is a differential operation and it inherits the blindness. The remedy is to have at least one absolute anchor in the design: one quantity whose value you know from outside the comparison, measured by an instrument you did not use to build either arm. Panel two has one, the satellite's own decommissioning date. Panel three's control does not, which is exactly why panel three can only rule out half of what it wanted to.


The check

Every number above is divided out in your browser from per-letter counts embedded in this page, and recomputed end to end, from the letters themselves, offline from the pinned source files by research/the-apparatus-is-the-finding/verify-the-apparatus-is-the-finding.mjs, which also string-matches the rendered figures out of this HTML and exits non-zero on any disagreement. What runs here, now:

The assertions below run here when the page loads, including the poison controls and a sweep of every setting the fault slider can reach. Every figure they use is printed in the prose above.

Recomputed in your browser.

Independent routes. Panel one's central claim is measured twice by instruments that share no code and no characters: a (b)(4) string counter, and a glyph counter written against a single letter and then run unchanged over all 458. They disagree about which batch is more redacted, by 5.19 times in opposite directions, and agree to 1.1 percent on the total. A third instrument, unrelated to either, counts the letterhead garble and the junk tokens and finds the two batches indistinguishable, which is what killed this page's first explanation of them. Across the page, the same structural prediction is put to three datasets whose analysis code has nothing in common: two confirm it and one refutes it, and the refutation is on the page.

Anchors nobody chose after seeing the answer. openFDA's own manifest publishes the record count and export date for this file, and the parse must reproduce them. FOIA has nine numbered exemptions, so a marker parse that emits an exemption number outside 1 to 9 is broken; the observed set is 1, 2, 3, 4, 5, 6, 7. Sentinel-1B's decommissioning on 23 December 2021 is stated in the source paper's own figure caption and predicts the sign and the location of the final-month collapse. Panel three has no anchor of that kind, and the check that looks like one is not: the GBIF facets were fetched with a year=1800,2026 filter, so their per-year buckets sum to the reported total by construction. That check catches a truncated file and nothing else, and it is labelled that way in the verifier rather than sold as an anchor here.

The poison control. The bench is the poison control, promoted to content. A fault of known size is applied to real data and the checker is required to catch it (arm A only) and then required to miss it (both arms) at the fault this page measured. Both directions are swept across all 201 slider settings rather than tested at one point, and the band where the both-arms battery stays quiet is printed above and asserted below. The offline verifier additionally poisons the panel-one letter texts and requires the batch ratio to move; reorders panel two's monthly changes, which preserves the endpoints and the whole multiset of changes and destroys the finding; and blends panel three's animal curve towards flat to find how far it would have to move before the control let this page's explanation live. Each restores and re-asserts clean afterwards.

Free choices, all of them. The glyph counter's character set was written from one letter and is listed above; two looser variants give a batch ratio of 0.97 and 1.01 instead of 1.011, in the same direction as that figure, and the verifier checks all three. The bench's anchor threshold is a factor of 1.5, declared before the result and shown with both values so you can see the margin. The bench split is deterministic (sort by length, alternate) and its permutation test uses 999 relabellings from a fixed seed. Panel three's decade sums use GBIF's basisOfRecord=PRESERVED_SPECIMEN facet; the unfiltered facet, pinned alongside it and checked by the same verifier, gives 19,633 and 6,853 for the 1930s and 2010s instead of 19,285 and 5,100, and does not change the shape or the control.

Named uncertainties. GBIF is live and grows; these facets are pinned to the fetch above and the most recent decade is the least stable. The openFDA export carries its own disclaimer that all results are unvalidated, and its approval_status column is a snapshot label, not an outcome: letters marked unapproved include live applications that may be approved later. Panel one shows that the two batches' texts were produced differently; it does not show which step differs, and it cannot, because this file holds no document that appears in both batches. The other class in the satellite file is unclassified detections, not a category of thing. Nothing on this page is a claim about whether any particular drug should have been approved, or about how much any agency redacts on purpose.