A combine portal · ten layers, one axis
Ten layers of this archive do the same honest thing: rebuild a published result under its authors' exact choices, enumerate every other choice that would have been defensible, and report where the published number sits inside the cloud that produces. Each one prints a percentile. Nobody had ever put those percentiles side by side, and doing it changes what they mean. The percentile is not a property of the finding. It is a property of the grid somebody chose to enumerate, and this page hands you the grid.
What is on this page
Every number below was read out of the ten member pages by booting each one in a real browser with the shared kit intercepted, so the values are the members' own, produced by their own shipped code on their own shipped data. The first thing this page checks about itself is that its percentiles are bit-identical to the ones those pages print.
One · the ruler
A specification curve answers a question of the form how far out on the range of defensible answers did the published one land? The answer is a percentile, and a percentile is comparable across subjects in a way the underlying quantities are not: an odds ratio, a warming trend, an earnings ratio and a language-split date share no units, but "sixty-ninth percentile of its own multiverse" means the same thing everywhere. So the first thing to do with ten such layers is to line them up.
Each row below is one published result. The filled dot is where its authors' own choice sits inside its own cloud. The pale ticks on the whisker are the same number ranked inside a grid one axis narrower, one tick per axis, computed live in your browser from the members' own specification values. Press the button to add the control.
the authors' own choice the same number, one axis held the same operation on a shuffled grid
Loading the fourteen grids…
Two · the counterfactual, by hand
Pick a result, then suppose its authors had never thought to vary one of their axes. The value in the cell does not change. The population it is ranked against does, and so the percentile does. This is not a criticism of any member: every one of them declares its grid in full. It is that the grid is a choice, made by one analyst, and the headline percentile inherits it.
| Layer | The result located | Defined cells | Own percentile | Held-axis range | Span |
|---|---|---|---|---|---|
| Loading… | |||||
Select a row above.
Three · the undeclared parameter
Every grid in this wave is a judgement about what counts as a defensible analysis, and every one of those judgements is an integer nobody argues about. The twelve grids these ten layers run span three orders of magnitude. Two of them evaluate fewer specifications than they declare, because the crossed product contains combinations the page rules out before running them, and the difference between those two numbers is not printed anywhere else.
| Layer | Axes | Declared | Run | Defined | Undefined |
|---|---|---|---|---|---|
| Loading… | |||||
Four · the control the wave contained and never used
A specification curve is a model of analyst behaviour. It says: here is what a room full of competent people might have done with this dataset. In exactly one member of this wave, that room was actually observed. Twenty-Nine Answers, One Dataset rebuilds Silberzahn and colleagues' crowdsourced study, in which twenty-nine independent teams analysed one football dataset and reported twenty-nine different answers, and it also enumerates a grid of 6,912 specifications over the same data.
That page asks whether the machine's range contains the humans. It does: twenty-eight of twenty-nine land inside. The question it does not ask is whether the machine is calibrated to them, and the answer is that it is not, in a way one number makes plain.
The obvious explanation is that the grid omits an axis the humans varied. It does omit one: seventeen of the twenty-nine teams describe a multilevel, hierarchical, mixed, clustered or Bayesian model, and no option anywhere in the grid's thirty-four declared levels is any of those. That explanation fails anyway. Those seventeen teams have a median odds ratio of — and the other twelve have —, which is the same number. The prediction was made here, tested here, and is reported as lost.
What is left is a property of the whole grid rather than of any axis in it. Slice the 6,031 defined specifications by every level of every axis, thirty-four slices in all, and take the median of each. Not one of the thirty-four reaches the median human team. The highest is —, still short of the humans' —. There is no axis you could drop, and none you could add from within the declared set, that closes the gap.
| Axis | Level | Cells | Median odds ratio | Distance below the median human team |
|---|---|---|---|---|
| Loading… | ||||
Five · what this does not say
Six · the check
The apparatus is research/the-ruler-nobody-declared/. It is three programs:
harvest.mjs boots each of the ten member pages in Chromium with
/_kit/multiverse.js intercepted and records every call the page makes into it;
derive.mjs reshapes that into the two binaries this page fetches;
verify.mjs re-derives every claim above and asserts it. Run it with
node research/the-ruler-nobody-declared/verify.mjs.
The load-bearing check is the first one. If the percentiles here were not bit-identical to the ones the member pages print, this portal would be a re-implementation quietly disagreeing with its sources, which is the most likely way a page like this lies. Three of the checks are controls designed to fail if the thing they guard is not really being measured, and two record predictions this work made and lost.