29 teams / 1 released file / 6,912 analysis choices

Twenty-Nine Answers, One Dataset

In 2018, 29 analyst teams answered the same question with the same football data. Their reported odds ratios ran from 0.89 to 2.93. This page rebuilds every published number, recomputes one raw-data model, then puts all 29 humans inside a machine grid that keeps its failed fits visible.

This is a page about analysis choices. The predictor is two raters' scores of profile photographs on a five-point light-to-dark scale. It is not race, ethnicity, or a referee's perception during play. No model here identifies a causal discrimination effect.

Loading the shipped data and running the checks


01 / Rebuild the published spread

First, make the anchor earn its place.

The paper prints a range of 0.89 to 2.93, a median of 1.31, and a 20 to 9 split between intervals above 1 and intervals crossing 1. Those values were reported in mixed units. The authors supplied the conversion arithmetic.

Authors' conversion, rerun in this browser

Waiting for 29 rows

d to OR: exp(d × π / √3)
r to d: 2r / √(1 − r²), then d to OR
OR and incidence-rate ratios: unchanged

Printed: 0.89 to 2.93, median 1.31, 20 / 9 / 0 significant positive / not significant / significant negative.

Reading the original-unit estimates now.

All 29 rounded values reproduce exactly. The residual difference from every printed two-decimal point is zero at the paper's precision. This is a reproduction of the published spread, not a pooled estimate. Silberzahn and colleagues did not report a pooled meta-analytic number.

The top of that range has been withdrawn by its own authors

verified at source, 2026-08-10

The range above is the one everybody quotes, and two of its largest values are ones the original authors no longer stand behind. Young and Stewart found that Team 21's Tobit estimate and Team 27's Poisson estimate came into line with the rest after routine adjustments. In a note posted to the study's own OSF project (osf.io/ch27w, Scaling.Issue.SilberzahnEtAl.pdf, October 2021), four of the original authors reply, in their words:

We (Silberzahn, Martin, Uhlmann, Nosek) agree that the revised estimates proposed by Y&S are sensible, and further comparably more defensible or “correct” than the estimates provided by the crowd analysts in Silberzahn et al. (2018).
The two revisions, as printed in that note
TeamOriginal methodOriginalRevised methodRevised
21Tobit regression2.88Fractional regression1.31
27Poisson regression2.93Poisson, non-significant interactions removed1.29

Recomputing the corrected spread from the same 29 rows.

The formal Corrigendum in the journal record (AMPPS 1(4):580) corrects a missing citation and changes no number, so a reader who follows the published literature alone will not meet this.

The authors also reject the conclusion you might expect us to draw, and their objection is aimed squarely at the kind of grid this page builds. They argue the correction “only underscore[s] the main arguments made in the original paper”, and that removing the two estimates entirely would not change their conclusions. Then they draw a distinction worth sitting with:

A crowd analysis captures the “research reality,” or actuarial question of what would happen if someone else analysed the data to test the same hypothesis. In contrast, a multiverse carries out all defensible specifications from the point-of-view of a single analyst or small team. … A crowd analysis is at once both less and more than a multiverse analysis.

That is a real limitation of everything below this line, stated by the people whose data it is. A grid enumerates what one mind judges defensible. It cannot enumerate what a stranger, with different training and different allegiances, would actually have done. Both numbers on this page are worth having, and neither replaces the other.


02 / One point from the raw rows

Before the grid, fit one ordinary specification.

Use every rated player, average the two photo ratings, count straight red cards, treat games as binomial exposure, and add no covariates. This point is not privileged. It is simple enough to audit.

...odds ratio
...model-based 95% CI
...inside clean grid
...released observations used

Change the uncertainty rule

The coefficient stays fixed. Only the interval changes. Player clustering aggregates exactly on the player table; referee clustering is recomputed from the shipped 124,621-dyad sidecar.

Computing interval

Exact sufficient-statistic check: the player-table and dyad-table tone coefficients differ by .... Aggregation is exact here because the predictor and every grid covariate vary only by player.


03 / Open every declared fork

The grid is 6,912 cells. It is not 6,912 successes.

Cross three units, six ways to use the two ratings, two card definitions, four model families, three exposure rules, and sixteen covariate sets. The slow fits took 464.5 seconds in the scout's fixed 25-step solver. To keep the page operable, the shipped cell estimates are precomputed. The browser builds the full factorial grid with the shared engine, joins every cell, recomputes all summaries, and refuses missing cells live.

Degraded, disclosed: this browser does not refit all 6,912 models. It refits the raw anchor, then loads the precomputed cell estimates. The verifier independently checks every displayed summary from those shipped cells, but does not rerun the 464.5-second fitting sweep.

6,031clean estimates summarized
870did not converge or were degenerate
11numeric separation, shown but quarantined

What the 870 cells were: every failure occurred in the cards-per-game or dyad-ever-carded units. There were 798 negative-binomial cells and 72 binomial cells. Poisson and linear-probability cells all returned. The 11 quarantined cells were also negative-binomial fits. Their returned OR was below 0.95 or above 4, including values from 0.001 to 5,155, under combinations with ignored exposure or unstable fixed effects. They remain a separate band rather than vanishing.

Joining the 6,912 cells

1.164clean median OR
1.092 to 1.250interquartile range
99.0%clean estimates above 1
35.7%model intervals above 1

Pick one cell

Any combination is reachable. If it did not return, the readout refuses the estimate and tells you why.

Building selected specification

Preparing the specification curve.

The plot window is fixed to OR 0.9 through 1.7 so the dense centre remains legible. The clean full range is 0.953 to 3.266. The summary uses all clean cells, not only the visible vertical window.


04 / Put the humans inside the machine

The range nearly encloses them. The density does not.

Auspurg and Bruederl already published a 486-model multiverse of this exact file: median OR 1.28, all estimates above 1, and 68% statistically significant. The contribution here is the live placement of all 29 human teams inside a wider declared grid, with its failures kept visible.

28 of 29published estimates inside the clean machine range
percentile 84.2published median inside the clean machine grid
2.2× widerhuman dispersion on the log-OR scale

The clean machine grid puts 99.0% of its mass above 1 and only 35.7% of model-based intervals wholly above 1. Direction and statistical significance are answering different questions. Neither number establishes causality.

Team 21 stays on the plot

2.88 printed
1.31 corrected

Young and Stewart showed that fractional regression removed the extreme value. The original authors accepted the correction publicly.

Team 27 stays on the plot

2.93 printed
1.29 corrected

Removing non-significant interactions removed the extreme value. The original authors accepted this correction too.

The printed values remain in the reproduction because that is the published anchor. The accepted corrections sit beside them because silently keeping the outliers would mislead.


05 / What moved the point

The answer depends on which grid survived.

On log odds ratios across the 6,031 clean cells, model family carries the largest main-effect share. Exposure rule is almost tied. Because 881 planned cells are absent from the clean summary, this decomposition is indicative, not an exact partition.

Main effects leave an interaction remainder of 55.9%. In the published 1,080-cell well-conditioned comparison, card definition dominated instead. The disagreement is not a contradiction. It is what happens when a different set of cells is admitted to the summary.


The check / limits before conclusions

What this instrument can and cannot say

Rating, not race

Two independent raters scored a profile photograph. Released values are already rescaled to 0, 0.25, 0.5, 0.75, and 1. One unit spans the whole light-to-dark scale.

1,585, not 1,586

The project README says 1,586 photographed players. The released file has 1,585 players with both ratings. This page uses and reports the released count.

Four leagues, one season

The rows cover England, France, Germany, and Spain in 2012 to 2013. This grid cannot generalise beyond that file.

No causal claim

Unobserved mediators remain plausible. Auspurg and Bruederl explicitly warn that the estimate is probably not a true causal effect.

Club is disputed

Some teams controlled for club. Auspurg and Bruederl argue club may be a collider. The grid keeps both branches rather than deciding that argument in advance.

Card definition is a fork

The data separate straight reds from second-yellow reds. Four one-game dyads also contain two yellow cards and no yellow-red, a quirk Team 1 flagged.

Slow fits are precomputed

The browser recomputes the 29-team anchor, the raw anchor, standard errors, grid summaries, controls, locations, and variance shares. It does not rerun the 464.5-second model sweep.

No null permutation

No permutation p-value is shown. A valid grid-wide shuffle would require refitting the models, which this degraded browser build does not pretend to do.


Sources and prior work

This grid has a history.