Team 21 stays on the plot
2.88 printed
1.31 corrected
Young and Stewart showed that fractional regression removed the extreme value. The original authors accepted the correction publicly.
29 teams / 1 released file / 6,912 analysis choices
In 2018, 29 analyst teams answered the same question with the same football data. Their reported odds ratios ran from 0.89 to 2.93. This page rebuilds every published number, recomputes one raw-data model, then puts all 29 humans inside a machine grid that keeps its failed fits visible.
This is a page about analysis choices. The predictor is two raters' scores of profile photographs on a five-point light-to-dark scale. It is not race, ethnicity, or a referee's perception during play. No model here identifies a causal discrimination effect.
Loading the shipped data and running the checks
01 / Rebuild the published spread
The paper prints a range of 0.89 to 2.93, a median of 1.31, and a 20 to 9 split between intervals above 1 and intervals crossing 1. Those values were reported in mixed units. The authors supplied the conversion arithmetic.
d to OR: exp(d × π / √3)
r to d: 2r / √(1 − r²), then d to OR
OR and incidence-rate ratios: unchanged
Printed: 0.89 to 2.93, median 1.31, 20 / 9 / 0 significant positive / not significant / significant negative.
Reading the original-unit estimates now.
All 29 rounded values reproduce exactly. The residual difference from every printed two-decimal point is zero at the paper's precision. This is a reproduction of the published spread, not a pooled estimate. Silberzahn and colleagues did not report a pooled meta-analytic number.
The range above is the one everybody quotes, and two of its largest values are ones the original authors no longer stand behind. Young and Stewart found that Team 21's Tobit estimate and Team 27's Poisson estimate came into line with the rest after routine adjustments. In a note posted to the study's own OSF project (osf.io/ch27w, Scaling.Issue.SilberzahnEtAl.pdf, October 2021), four of the original authors reply, in their words:
We (Silberzahn, Martin, Uhlmann, Nosek) agree that the revised estimates proposed by Y&S are sensible, and further comparably more defensible or “correct” than the estimates provided by the crowd analysts in Silberzahn et al. (2018).
| Team | Original method | Original | Revised method | Revised |
|---|---|---|---|---|
| 21 | Tobit regression | 2.88 | Fractional regression | 1.31 |
| 27 | Poisson regression | 2.93 | Poisson, non-significant interactions removed | 1.29 |
Recomputing the corrected spread from the same 29 rows.
The formal Corrigendum in the journal record (AMPPS 1(4):580) corrects a missing citation and changes no number, so a reader who follows the published literature alone will not meet this.
The authors also reject the conclusion you might expect us to draw, and their objection is aimed squarely at the kind of grid this page builds. They argue the correction “only underscore[s] the main arguments made in the original paper”, and that removing the two estimates entirely would not change their conclusions. Then they draw a distinction worth sitting with:
A crowd analysis captures the “research reality,” or actuarial question of what would happen if someone else analysed the data to test the same hypothesis. In contrast, a multiverse carries out all defensible specifications from the point-of-view of a single analyst or small team. … A crowd analysis is at once both less and more than a multiverse analysis.
That is a real limitation of everything below this line, stated by the people whose data it is. A grid enumerates what one mind judges defensible. It cannot enumerate what a stranger, with different training and different allegiances, would actually have done. Both numbers on this page are worth having, and neither replaces the other.
02 / One point from the raw rows
Use every rated player, average the two photo ratings, count straight red cards, treat games as binomial exposure, and add no covariates. This point is not privileged. It is simple enough to audit.
The coefficient stays fixed. Only the interval changes. Player clustering aggregates exactly on the player table; referee clustering is recomputed from the shipped 124,621-dyad sidecar.
Exact sufficient-statistic check: the player-table and dyad-table tone coefficients differ by .... Aggregation is exact here because the predictor and every grid covariate vary only by player.
03 / Open every declared fork
Cross three units, six ways to use the two ratings, two card definitions, four model families, three exposure rules, and sixteen covariate sets. The slow fits took 464.5 seconds in the scout's fixed 25-step solver. To keep the page operable, the shipped cell estimates are precomputed. The browser builds the full factorial grid with the shared engine, joins every cell, recomputes all summaries, and refuses missing cells live.
Degraded, disclosed: this browser does not refit all 6,912 models. It refits the raw anchor, then loads the precomputed cell estimates. The verifier independently checks every displayed summary from those shipped cells, but does not rerun the 464.5-second fitting sweep.
What the 870 cells were: every failure occurred in the cards-per-game or dyad-ever-carded units. There were 798 negative-binomial cells and 72 binomial cells. Poisson and linear-probability cells all returned. The 11 quarantined cells were also negative-binomial fits. Their returned OR was below 0.95 or above 4, including values from 0.001 to 5,155, under combinations with ignored exposure or unstable fixed effects. They remain a separate band rather than vanishing.
Joining the 6,912 cells
Any combination is reachable. If it did not return, the readout refuses the estimate and tells you why.
Preparing the specification curve.
The plot window is fixed to OR 0.9 through 1.7 so the dense centre remains legible. The clean full range is 0.953 to 3.266. The summary uses all clean cells, not only the visible vertical window.
04 / Put the humans inside the machine
Auspurg and Bruederl already published a 486-model multiverse of this exact file: median OR 1.28, all estimates above 1, and 68% statistically significant. The contribution here is the live placement of all 29 human teams inside a wider declared grid, with its failures kept visible.
The clean machine grid puts 99.0% of its mass above 1 and only 35.7% of model-based intervals wholly above 1. Direction and statistical significance are answering different questions. Neither number establishes causality.
2.88 printed
1.31 corrected
Young and Stewart showed that fractional regression removed the extreme value. The original authors accepted the correction publicly.
2.93 printed
1.29 corrected
Removing non-significant interactions removed the extreme value. The original authors accepted this correction too.
The printed values remain in the reproduction because that is the published anchor. The accepted corrections sit beside them because silently keeping the outliers would mislead.
05 / What moved the point
On log odds ratios across the 6,031 clean cells, model family carries the largest main-effect share. Exposure rule is almost tied. Because 881 planned cells are absent from the clean summary, this decomposition is indicative, not an exact partition.
Main effects leave an interaction remainder of 55.9%. In the published 1,080-cell well-conditioned comparison, card definition dominated instead. The disagreement is not a contradiction. It is what happens when a different set of cells is admitted to the summary.
The check / limits before conclusions
Two independent raters scored a profile photograph. Released values are already rescaled to 0, 0.25, 0.5, 0.75, and 1. One unit spans the whole light-to-dark scale.
The project README says 1,586 photographed players. The released file has 1,585 players with both ratings. This page uses and reports the released count.
The rows cover England, France, Germany, and Spain in 2012 to 2013. This grid cannot generalise beyond that file.
Unobserved mediators remain plausible. Auspurg and Bruederl explicitly warn that the estimate is probably not a true causal effect.
Some teams controlled for club. Auspurg and Bruederl argue club may be a collider. The grid keeps both branches rather than deciding that argument in advance.
The data separate straight reds from second-yellow reds. Four one-game dyads also contain two yellow cards and no yellow-red, a quirk Team 1 flagged.
The browser recomputes the 29-team anchor, the raw anchor, standard errors, grid summaries, controls, locations, and variance shares. It does not rerun the 464.5-second model sweep.
No permutation p-value is shown. A valid grid-wide shuffle would require refitting the models, which this degraded browser build does not pretend to do.
Sources and prior work