Afterthoughts / Strata Evidence firewall: 31 Jul 2010

Control 3 / Ego depletion / Historical replay

The warning in plain sight

The 2010 meta-analysis said the effect was large and robust. Its own archive, checked with methods already in print, supported a more cautious decision: run one large, preregistered replication before trusting the headline.

2010 prose d = .62

A seemingly decisive average from 198 effects and 10,782 participants.

Rule frozen to 2010 tools Warn

Precision dependence plus a material correction triggers a decisive replication.

2016 registered replication d = .04

23 laboratories, N = 2,141, 95% CI [-.07, .15]. The decision was right. The corrections were not equally predictive.

The check

Computed live from 198 rows
Annual endpoints only. The source does not preserve within-year release order.
Outlier values

This changes the highlighted estimate, never the underlying rows.

Effects / participants198 / 10,782
Ordinary fixed d.623
Trim-and-fill d.477
PET intercept-.095

Contour-enhanced funnel. Dashed lines mark two-sided p = .05 around zero. Imputed points are outlined.

Used198
Later0
Excluded0
Undated0
Total198
Uncertainty and free choices

The shipped file is the 198-effect data sheet reproduced by Blázquez, Botella, and Suero (2017), which in turn reproduces the Hagger et al. archive. The machine-readable Dryad deposit was released in 2015, not 2010; it is evidence about the 2010 information set, not a file a 2010 analyst could download. Study labels supply publication year only, so rows retain that coarse year timestamp and the replay stops at annual endpoints. The label “in press” is normalized to the final 2009 publication year. The default preserves Hagger’s three adjusted outliers. The raw switch restores 3.02, 2.60, and -0.57. Sampling variance is (n1+n2)/(n1 n2) + d²/[2(n1+n2)]. Regressions use inverse-variance weights. The 1.96 intervals are large-sample intervals. Bias procedures are sensitivity analyses, not proof that missing studies exist.

Anchor reconciliation

Firewall negative control

Frozen before viewing 2016

Demand a decisive replication when all three conditions hold.

  1. At least 20 effects are available.
  2. Egger’s precision-dependence slope has p < .10.
  3. Either trim-and-fill moves d by at least .10 or PET’s 95% interval includes zero.

This is a decision rule, not an estimator. Every ingredient was published before the July 2010 cutoff.

Complete annual replay

The rule first warns in 2003. It never reverses.

The replay runs through the entire archive. No convenient stopping year is omitted. Early years with fewer than 20 effects are automatically ineligible.

YearkOrdinary dTrim-fill dPET dDecision

Historically admissible methods

Every applied bias check existed by 2010.

The date is part of the method. These are the only bias-detection or correction procedures that enter the historical reconstruction.

Egger regression Applied

Tests whether effect size depends on its standard error. Here, the inverse-variance regression is d = β₀ + β₁SE; the test is β₁ = 0.

Egger, Smith, Schneider, and Minder, “Bias in meta-analysis detected by a simple, graphical test,” BMJ 315:629-634, published September 13, 1997.

Trim-and-fill Applied

Uses Duval and Tweedie’s L0 rank estimator, trims the asymmetric right tail, then mirrors effects around the refitted center.

Duval and Tweedie, “A nonparametric ‘trim and fill’ method of accounting for publication bias in meta-analysis,” Journal of the American Statistical Association 95:89-98, published March 2000; algorithm also in Biometrics 56:455-463, June 2000.

PET regression Applied

The intercept from inverse-variance weighted d on SE estimates the effect as SE approaches zero. PET is shown alone, with no later conditional selection rule.

Stanley, “Meta-regression methods for detecting and estimating empirical effects in the presence of publication selection,” Oxford Bulletin of Economics and Statistics 70:103-127, first published online September 19, 2007; 2008 issue.

PEESE variance regression Applied

The intercept from inverse-variance weighted d on sampling variance is displayed as a separate sensitivity estimate.

Moreno et al., “Assessment of regression-based methods to adjust for publication bias through a comprehensive simulation study,” BMC Medical Research Methodology 9:2, published January 9, 2009.

Excess significance Applied

Counts significant effects and compares them with the count expected from power under the common-effect estimate. It inherits that estimate as a free choice.

Ioannidis and Trikalinos, “An exploratory test for an excess of significant findings,” Clinical Trials 4:245-253, published June 2007.

Contour-enhanced funnel Applied

Adds the p = .05 contours to the funnel so the locations of the L0 imputations are visible. Asymmetry can have causes other than publication bias.

Peters, Sutton, Jones, Abrams, and Rushton, “Contour-enhanced meta-analysis funnel plots help distinguish publication bias from other causes of asymmetry,” Journal of Clinical Epidemiology 61:991-996, published October 2008.

Control 3

The warning wins. The corrections do not.

A frozen warning rule only had to say “do the decisive test.” It did. Asking each correction to forecast the later effect is harder, and the honest scorecard is mixed.

Trim-and-fill.48

Still far above the 2016 estimate of .04.

PEESE, raw sensitivity.25

Lower, but its 95% interval [.18, .32] still misses .04.

PET, raw sensitivity-.11

Its published 95% interval [-.23, .02] includes .04 and zero. It predicted the later result better.

Values shown above use restored raw outliers where needed to reproduce Carter and McCullough’s 2014 regression sensitivities; the live default preserves Hagger’s 2010 adjustments. The point is not that PET is a time machine. It is that the archive justified a warning even though its corrections disagreed.

POST-CUTOFF: NOT USED BY THE 2010 DECISION RULE

Later methods

Useful now. Forbidden in the 2010 decision.

Conditional PET-PEESE

The rule that selects PEESE after a significant PET result is post-cutoff. Stanley and Doucouliagos’s paper appeared online November 20, 2013 and in Research Synthesis Methods 5:60-78 in March 2014. It is not applied above.

P-curve

P-curve first appeared online in 2013 and in Journal of Experimental Psychology: General 143:534-547 in 2014. It is not applied anywhere on this page.

Anytime-valid e-process

The monitoring audit at right is present-day site machinery, implemented in 2026. It checks the danger of repeatedly peeking at accumulating evidence. It is neither a 2010 bias detector nor an input to the frozen warning rule.

Peeking calibration, seeded and live

Under a Bernoulli null with p = .50, 1,000 deterministic simulated streams of length 198 are watched after every record. The naive one-sided fixed-sample p-value is flagged whenever it falls below .05. The mixture e-process is flagged at 20, the Ville threshold for α = .05.

Naive p ever crossescomputing
E-process ever crosses 20computing
Ever-crossing rate as peeks accumulate
PeeksNaive p < .05e ≥ 20

Actual stream calculation pending.

Within-year release order is not recoverable. The actual stream uses publication year, then source-file order, and is interpreted only at annual endpoints. Seed: 20100731. This panel demonstrates monitoring calibration; it does not estimate ego depletion.

Control 4 / The unavailable quantity

Nothing in the 2010 archive contained the 2016 answer.

No amount of diligence in July 2010 could reveal the effect under one shared, preregistered protocol across 23 laboratories: d = .04 with interval [-.07, .15]. Those observations did not exist.

What the record could support

It could support a decision to commission a decisive replication. It could reveal precision dependence, funnel asymmetry, excess significance, and strong sensitivity to the chosen correction.

What it could not support

It could not identify which bias correction would land closest to a future protocol-constrained estimate. It could not separate publication selection from every other cause of small-study effects. It could not tell us how the literature or theory would have changed had a registered replication been demanded in 2010.

The counterfactual stops at the decision: the frozen rule would have asked for the later kind of test. This page scores that rule against the outcome that actually arrived. It does not claim knowledge of the branch history did not take.

Provenance

Sources and audit trail

  1. Hagger et al. (2010), Psychological Bulletin. The published meta-analysis and 198-effect anchor.
  2. Hagger et al. Dryad deposit. Machine-readable supporting data, released March 24, 2015 under CC0.
  3. Blázquez, Botella, and Suero (2017). CC BY supplementary data sheet used on this page, containing the same 198 effects.
  4. Hagger et al. (2016), Perspectives on Psychological Science. Registered replication report, 23 laboratories and N = 2,141.
  5. Carter and McCullough (2014), Frontiers in Psychology. Published sensitivity targets for trim-and-fill, PET, PEESE, and excess significance.
  6. Egger et al. (1997); Duval and Tweedie (2000); Stanley (2008); Moreno et al. (2009). Historical method sources.

The exact shipped computations and adversarial checks are rerunnable with node research/ego-depletion-warning-in-plain-sight/verify.mjs. No result on this page requires a network request.