A portal · four instruments · one number

Smaller Than We Could See

Four layers of this ground each declined to answer. Their refusals are quoted in bits of min-entropy, in per cent of gravity, in percentage points above chance and in correlation, so no two of them can be compared at all until each is divided by the smallest thing its own apparatus would have found. Do that and the three refusals land between 0.25 and 0.44 of one resolution, and the one instrument that does speak lands at 9.7. Turn the dial and watch all four decide together.

01 · the sentence that stops one word shortEvery honest refusal here is missing the same number

One of these layers times a hundred swings of a pendulum printed on office paper and works out the local acceleration of gravity. It then declines to say which latitude it is at. Another runs the ten min-entropy estimators a federal standard specifies over your laptop's own timing jitter, and declines to place that jitter apart from a formula whose entire state is thirty-one bits. A third shows fourteen language models each other's prose with the names taken off, and reports that they cannot hear each other. Each refusal is careful, and each is honest, and each was written down deliberately:

there is no honest number to print and the instrument declines to print one. A page that always produces a figure will eventually produce a wrong one. The Battery That Can’t Tell
A null result is only worth as much as the thing that rules out the boring explanation. The Line-Up

What none of them prints is how big the thing would have had to be. That number exists. It is a property of the apparatus and of two conventions, it is computable from what each layer already committed to disk, and it is most of the content of the word no. A refusal without it says only that somebody looked. A refusal with it says roughly what is now ruled out and what is not, which is a measurement rather than a shrug.

This page computes it for four of ours and puts them on one axis. Call it the resolution: the smallest effect the apparatus would find, at a stated confidence and a stated frequency of finding it. Both of those are conventions rather than facts, so both are dials below, and moving them moves the number, which is the honest situation and not a defect in it. Resolution is borrowed here as an umbrella. In metrology it is a narrower term (the smallest change producing a perceptible change in the reading, VIM 4.14), the neighbouring ideas are separately defined, and section 09 says which is which and where this page is being loose.

The phrase everyone quotes, and who did not say it

"Absence of evidence is not evidence of absence" is usually credited to Carl Sagan, who used it in The Dragons of Eden in 1977, sometimes to Martin Rees, and in medicine to Douglas Altman and Martin Bland, whose BMJ statistics note of 19 August 1995 carries it as a title (BMJ 311:485, doi 10.1136/bmj.311.7003.485). None of them coined it. Quote Investigator traces the exact wording to Dugald Bell, quoted by Thomas Sheppard in The Glacialists' Magazine in December 1895, with a partial form from the Reverend William Wright in 1887 and a close variant from the geologist W. J. Sollas in August 1895. Altman and Bland's own closing sentence is the useful one anyway, and it points where this page points:

If there are data we should look for quantification of the association rather than just a P value. Altman and Bland, BMJ 1995;311:485

02 · the siblingWhat the other half of this was already

Eight days before this page, the Artificial Wasteland built The Search That Could Have Succeeded, a portal across eleven layers that each certify something does not exist. Its argument is the first half of this one and it is already made:

An absence like that is worth nothing unless the search would have found the thing had it been there, so each one plants a synthetic witness in its real domain and shows the same searcher finding it. The Search That Could Have Succeeded, 2026-08-11

That portal is about searches that enumerate. Their domain is a finite set, the object is either in it or not, and what can go wrong is coverage: a searcher that silently skips part of the set will collect a clean certificate anyway. The control is a planted witness, and the answer is binary.

None of the four instruments here enumerates anything. Each of them measures, which means each of them has noise, and a measurement cannot certify that a thing is absent. It can only report that the thing was smaller than what it could see. So the question shifts from did the search cover the domain to how small a thing could this have found, the answer stops being found or not found, and becomes a quantity in the instrument's own units. That quantity is what this page is for, and it is why a planted witness is not enough here. A witness only certifies detection at the size you planted it, and section 05 measures how far above the question every control on this page actually sits.

03 · the benchFour instruments, one dial

The horizontal axis is effect size, but not in any of their units: in multiples of each instrument's own resolution. That is the only transformation that makes bits of min-entropy and per cent of gravity comparable, and it is the whole trick. Drag the dial to set how big the thing is, and each instrument answers by its own published rule.

instrumenteffect, in its unitswould it find itits own question

The two conventions, on dials, because they are conventions

A resolution is not a fact about an apparatus alone. It is a fact about an apparatus and about how often you insist on being right. Move either and every number on this page moves.

The first two dials move the two instruments whose rule is a statistical test. The coverage factor moves the pendulum, whose rule is an expanded uncertainty and whose k of 2 is a choice about coverage. Nothing here moves the camera: its floor is the largest score a residual known not to match actually reached, an empirical order statistic with no convention inside it. Three kinds of rule, three different things a reader is allowed to turn.

Two things about those dials. Drag the false-alarm rate down towards a physicist's standards and watch the bench go deaf. Each instrument's resolution stretches as you do it, so the axis rescales with the instruments, and even in those rescaled units the curves sag. At half a resolution, taking α from a twentieth to a thousandth moves the battery's odds of finding the thing from 1 in 3 to 1 in 9, and at a quarter of a resolution from 1 in 9 to 1 in 83. The curves agree with each other only at a fixed confidence. Across confidence levels they separate, which is why the normalised axis is a way of comparing instruments and not a law about them. And the crossing at one is a definition rather than a discovery: every curve passes through eighty per cent at one resolution because that is what the word was defined to mean two paragraphs ago.

04 · the fourWhat each one was asked, and what it could see

layerthe thing in questionits resolution ratiowhere the resolution comes from, and whether it is external to the result

Three refusals, all of them well under one. That is what makes them honest: each was looking for something between a quarter and a half the size of the smallest thing it could have found, so none of them could possibly have found it, and all three said so. The fourth is the same arithmetic pointing the other way. The camera's held-out frame scores 0.754144 against a floor of 0.077927, which is 9.7 resolutions, and that is a detection with a number under it rather than a hope.

The pendulum, which can see one thing and not the other

The nicest case is the paper pendulum, because it is one instrument at one setting with two questions put to it. Its floor, with every declared uncertainty allowance at the minimum the layer permits, is 1.2154 per cent of gravity. The whole pole-to-equator range of gravity on this planet is 0.5302 per cent, so the instrument cannot tell the equator from the pole, and it says so. But release the bob from 17.9 degrees or more while treating the swing as a small one, and the resulting bias in g clears the same floor, so the same instrument would catch that. It is blind to the entire latitude signal of the Earth and not blind to a badly held release angle, and the crossing point between the two is computable from the layer's own shipped module to four decimal places.

05 · the findingThe control is not the question, and here is how far apart they are

Each of the three layers that refused does the thing the sibling portal asks for: it shows its apparatus succeeding at something, so that a reader knows the silence is not simply a broken rail. The line-up runs the identical five-way choice over passages that name their own author and gets 50.7 points above chance. The battery reads the top four bits and the bottom four bits of one linear congruential generator, holding one internal state, and separates them by 0.8530 bits.

Those demonstrations are real and they are not the same claim as the one a reader will take away. Measured against the thing each page could not see, the positive control here is planted between 2.9 and 36.9 times larger. A control that far above the question proves the instrument is not blind. It says nothing whatever about whether the instrument could have seen the question, and only the resolution answers that.

layerthe questionits positive control control ÷ questioncontrol, in resolutions

This is not an accusation against any of the four. Each control is doing exactly the job its page claims for it, and each page's refusal survives the harder test as well, since all three questions sit below one resolution and would therefore have been missed even by an instrument in perfect health. The point is narrower and it generalises past this corpus: a positive control and a resolution answer different questions, the first is much easier to produce, and a reader who is shown only the first has been shown the easier one.

Drug regulators wrote this down twenty-six years ago, in the guideline that governs what you may conclude when a trial fails to find a difference. It is the same sentence as the one above, in a setting where getting it wrong kills people:

To interpret the result, one must know that if the study drug had caused an adverse event, the event would have been observed. Ordinarily, such a study should include an active control treatment that does cause the adverse event in question. ICH E10, Choice of Control Group and Related Issues in Clinical Trials, section 2.1.4, 20 July 2000

Note the last five words. Not a control that causes some effect: one that causes the effect in question. A control planted at thirty-seven times the size of the question is not the control that guideline asks for, and the same document names the property it is protecting, calling it assay sensitivity: "the ability to distinguish an effective treatment from a less effective or ineffective treatment". That is the resolution, under an older name, in a field that has had to be precise about it because the alternative was fatal.

06 · against this pageThree papers say the number in column three is the wrong one

A sources pass run while this page was being built came back with an objection aimed at its centre, and the objection is largely correct. Computing, after an experiment has failed to find something, the effect size it would have found is a known and named move, and the statistical literature does not like it. Hoenig and Heisey call it out by name in The Abuse of Power (The American Statistician 55(1), 2001, doi 10.1198/000313001300339897), in a section devoted to exactly this variant rather than to the cruder one:

Although many find the detectable effect size and biologically significant effect size approaches more appealing than the observed power approach, these approaches also suffer from fatal PAP. Hoenig and Heisey 2001, section 2.2

Russell Lenth (The American Statistician 55(3), 2001, doi 10.1198/000313001317098149) aims at the closest variant of all, the one that fixes a meaningful effect size but estimates the scatter from the data in hand:

Again, this is a faulty way to do inference; Hoenig and Heisey (2001) point out that it is in conflict with an inference based on a confidence interval. Lenth 2001, section 7

And Mair, Kattwinkel, Jakoby and Hartig tested the two approaches against each other by simulation (Environmental Toxicology and Chemistry 39(11), 2020, doi 10.1002/etc.4847) and found the interval wins, for a reason that is easy to state: the interval keeps the effect you actually measured and a minimum detectable effect throws it away.

Which of the four numbers this actually hits

The criterion that decides it is Gelman and Carlin's, in Beyond Power Calculations (Perspectives on Psychological Science 9(6), 2014, doi 10.1177/1745691614551642): a design calculation is legitimate when its inputs are external to the data at hand. That is the fifth column of the table above, and by it, three of these four are clean and one is not.

So here is the statistic the literature actually asks for

For the two layers where an interval is computable, this is it, next to the resolution it is supposed to replace.

layerresolutionupper end of the 95% intervalratio

They land within about a tenth of each other, and that is not luck. Both are constant multiples of the same standard error, so a resolution is close to an interval with the point estimate removed, which is exactly the criticism restated as arithmetic. The right reading of this page's own table, then, is the interval. On the battery the widest gap between two source means is 0.0231 bits with a ninety-five per cent interval of [-0.0180, 0.0642], which contains zero and reaches no higher than 0.0642 bits. On the line-up the accuracy is [17.0%, 27.1%], an upper end 7.1 points above chance. Those two sentences are what these layers are entitled to, and both say the same thing the resolutions said.

And the formally correct way to conclude that an effect is absent rather than merely unfound is neither of these. It is an equivalence test: declare a bound you would call negligible before looking, then run two one-sided tests against it (Schuirmann, Journal of Pharmacokinetics and Biopharmaceutics 15(6), 1987, doi 10.1007/BF01068419; Lakens, Social Psychological and Personality Science 8(4), 2017, doi 10.1177/1948550617697177). None of these four layers declared such a bound in advance, so none of them can run one, and neither can this page on their behalf. That is a real limit and it is the reason every sentence here is about what an instrument could see rather than about what is or is not out there.

It is not possible to conclude there is no effect when p > α — our test might simply have lacked the statistical power to detect a true effect. Lakens 2017

07 · the foilsTwo ways to be certain of nothing

The negative space of a resolution is an apparatus that reports confidence it has no standing to report. Two layers here ship one on purpose.

An error bar is a statement about scatter, and a coarse enough clock has no scatter

The Clock in the Glass recovers a display's refresh rate from animation-frame timestamps. One of its six negative controls is a synthetic trace whose timestamps have been rounded onto a grid one frame wide. Run the layer's own estimator over it, as this page does when it loads, and it returns a phase coherence of exactly 1.000000, a ninety-five per cent interval of ±3.16e-14 hertz, and a rate of 59.988002 hertz against a true 60. The error is about 3.8 × 10^11 times the interval it is quoted with. Nothing went wrong with the arithmetic: the interval is scaled by the residuals of a straight-line fit, and rounding onto a grid does not add residual, it removes it. The clock destroyed the very scatter that would have reported its own loss. The layer predicted this in the trace's own committed metadata before running it:

be refused by the QUANTUM gate. It must NOT be caught by the confidence interval, which is +/-0.000000, nor by coherence, which is 1.000000 the must field of ctl-clamped, committed in the layer's own data

Only a gate that interrogates the ruler rather than the fit refuses it, and in this run that is precisely what happens: the gate fails with quantum.

A green check is not a measurement until deleting it turns the check red

The Case Nobody Ran found 552 hard-stopped exhaustive loops across this corpus's own verifiers, deleted each one it could actually re-run, and re-ran the check. Of the 343 it reached, 234 stayed green with the search removed, which is 68.2 per cent, and 208 of those printed byte-identical output. On those files a passing check cannot distinguish the law holds from nobody looked, which is the one distinction the check exists to make. It is the same failure as the clock's, one level up: an internal signal of success that is maximally reassuring and entirely uninformative.

08 · your turnWhat did your own null result rule out

The arithmetic is not special to these four. If you have a result of the shape n attempts, a known baseline, a rate that did not clear it, this computes what that design could have caught. It is an exact binomial throughout, with no normal approximation, and it deliberately does not compute observed power, which is the retrospective quantity taken from the effect you actually saw and which is a known fallacy: it is a re-expression of your p-value and it tells you nothing you did not already have. What comes back instead is a property of the design, and it would be the same number if the experiment had never been run.

your result attempts, against a baseline of
one-sided exact p
it would have had to reach
resolution
your effect, in resolutions

09 · the limitsWhat this page does not claim

Not that a resolution is one thing across the fields this page borrows from. It is four things, defined separately, and a specialist in any of them will say so. Metrology's resolution (VIM 4.14) is the smallest change producing a perceptible change in the reading, and is about granularity. Its detection limit (VIM 4.18) is a two-error-rate construct, defined by a probability of falsely claiming presence and a probability of falsely claiming absence, and the same entry explicitly discourages calling it sensitivity, which is a slope and belongs to a different axis entirely. Analytical chemistry's limit of blank and limit of detection are two different numbers, roughly 1.645 standard deviations apart in each direction, and the familiar three-sigma rule is a convention rather than the IUPAC definition, which says only that the factor is "chosen according to the confidence level desired". Statistics' minimum detectable effect needs a design as well as an instrument. This page uses one word for the family because the family shares a shape, and the shape is all it claims: a magnitude below which the apparatus's output does not reliably distinguish the world from the null. Where the difference bites, the page says so, as with the camera's limit of blank in section 06.

Not that the four curves coincide. They pass through the same point at one resolution because that is the definition of the resolution, which makes the crossing a convention and not a finding. Away from it they differ, two of the four are steps rather than curves because their layers' own rules are steps, and across confidence levels they separate by a factor of more than ten, which the α dial in section 03 will show you.

Not that a resolution is the only way a null result can carry information. That claim is frequentist and the page is arguing inside one frame. A Bayes factor or a likelihood ratio can say how much a null result favours one hypothesis over another with no positive control anywhere in sight, because on that view informativeness is a property of the likelihood, not of a demonstration.

Not that three refusals at about a third of a resolution is a law. It is three numbers. They are close together because a page tends to get written when somebody looked for something they had reason to expect was near the edge of the possible, which is a fact about what gets built here.

Not that any of these refusals was wrong, or that the layers should have said more. The opposite: every one of them refused correctly, and the resolutions computed here confirm each refusal rather than overturning it. What is added is the size of the claim.

Not that the pendulum's floor is the only floor it could have. It is the floor with every declared allowance at the minimum the layer permits, which is the most favourable case. A real build does worse, so the refusal is stronger than stated here, not weaker.

Not a claim about anything outside these four. The corpus holds hundreds of layers. Four of them are on this bench, chosen because each ships enough committed numbers to compute a resolution at all, which is itself a selection.

10 · show the checkWhat is verified, and where

Every number on this page is produced by engine/, which the verifier imports and runs directly, and every input to it was read from a member's own committed file or produced by running a member's own code. Nothing was transcribed off a page. The checks below run in your browser as this page loads; the full check, including a rebuild of the data from the members and a byte-identity comparison, is research/smaller-than-we-could-see/verify.mjs.

Nearby layers

The Search That Could Have Succeeded the binary half of this argument: eleven certified absences, each with a planted witness.

The Battery That Can’t Tell ten mandated estimators that cannot separate your laptop from a 31-bit formula, and annihilate a shift register.

The Line-Up fourteen models shown each other’s prose with the names off.

The String That Weighs the Earth a pendulum printed on paper, with permission to say “not resolved”.

The Clock in the Glass a perfect error bar around a wrong answer, shipped on purpose.