At full strength / psychology / a replication in the open
Thirty-Seven Thousand Curtains
In 2011 Daryl Bem reported that people chose the curtain that would hide an erotic picture 53.1% of the time, although the computer picked its side only after they had chosen. Recompute that from his own session data, then open the 37,836 curtains of the preregistered replication he helped design, which found 49.89%, and plant his 53.1% into a copy of those trials to see whether the replication could have caught it.
Everything below computes in your browser from four small files you can read: Bem's one hundred sessions, and every erotic trial the replication recorded, in the order its stopping rule counted them. Numbers that come from a published paper are marked as printed and cited where they appear. Nothing is uploaded.
jumps to the replication and replays it; the claim it tests comes first, just below.
I / The claim, at the strength it was printed
Pick a curtain. Then the computer decides where the picture goes.
In 2011 the Journal of Personality and Social Psychology published nine experiments by Daryl J. Bem, of Cornell University, reporting that people's responses were influenced by random events that had not yet happened. The first of them, run with 100 Cornell undergraduates, is the one above. Here is its result, in his words.
Across all 100 sessions, participants correctly identified the future position of the erotic pictures significantly more frequently than the 50% hit rate expected by chance: 53.1%, t(99) = 2.51, p = .01, d = 0.25.Bem (2011), Journal of Personality and Social Psychology 100(3), p. 409 (version of record). Significance levels in the article are one-tailed and d is the effect size (footnote 3).
And the claim was a programme, not a single number. The abstract of the same article:
The mean effect size (d) in psi performance across all 9 experiments was 0.22, and all but one of the experiments yielded statistically significant results.Bem (2011), abstract, version of record.
Bem's 100 sessions, recomputed here from his own data
Pooled over trials: 826 hits in 1,560 erotic trials (52.95%), z = 2.30 with a continuity correction, exact one-tailed p = 0.0106 (printed z = 2.30, p = .011). Nonerotic trials: 49.83%, t = −0.15 (printed 49.8%, t = −0.15). Erotic minus nonerotic: t = 1.86, d = 0.19 (printed 1.85 and 0.19; the unrounded value is 1.8563, a gap in the last printed digit).
The engine has not run yet in this browser.
Every one of those numbers comes out of one function, the session t test Bem describes, run on a per-session extract of the workbook he released in 2018 (hits and trials per session, derived from its percentages). The claim reproduces at full strength: the page's engine sees the effect exactly as its author printed it.
All 100 sessions: 53.14%, t = 2.51.
The Method section describes the two halves as two sets of sessions, and this page did not choose the
split (section VII gives a published objection to it and Bem's reply). From the Method section: 40 of the sessions comprised 12 trials using erotic pictures, 12 trials using negative pictures, and 12 trials using neutral pictures
, with sides
drawn by the programming language's own random function, and the remaining 60 sessions comprised 18 trials using erotic pictures and 18 trials using nonerotic positive pictures
, with sides drawn by a hardware random number
generator (p. 409). Sessions 1 to 40 average 54.37% (t = 2.19); sessions 41 to 100 average
52.31% (t = 1.44). The workbook agrees independently: 37 of the first 40 sessions
cannot be read as 18 and 18 trials, and 56 of the last 60 cannot be read as 12 and 24, because their
percentages would not be whole numbers of hits. The anchor always uses all 100.
826, not 828. The workbook holds 826 erotic hits in 1,560 trials
(52.95%). The printed 53.1% is the mean of the 100 session percentages (53.139%); the binomial z = 2.30 reproduces
from 826. The replication used 828 of 1,560, which is 53.1% of 1,560 rounded, and printed Bem reported 53.07%
. This
page uses 828 wherever it reproduces the replication and 826 wherever it computes from Bem's data.
II / The deciding control
Open the curtains the replication opened.
The Transparent Psi Project repeated Experiment 1 in ten laboratories in nine countries between 10 January 2020 and
29 April 2022, under a protocol fixed in advance by a panel that included the claimant: our study protocol was based on a consensus of a panel of experts, which involved Daryl Bem
. It
was published in Royal Society Open Science on 1 February 2023. An earlier preregistered replication, by Wagenmakers and
colleagues, had also found nothing, but in the replication paper's words it was run independently of Bem, who later criticized the protocol for using an inadequate stimulus set
.
This one was designed with him.
It did not copy every detail. The paper prints its deviations: we used a different set of erotic images than in the original study for legal reasons
; multiple participants could be tested in the same experimental space at the same time
;
payment differed between sites. All of these deviations were accepted by the consensus panel.
The rule, as registered
The decision was not one test. At each of five looks (after 37,836, 62,388, 86,958, 111,528 and 136,080 erotic trials, counted in the order they reached the database) four tests run:
- a logistic regression with a random intercept for each participant, whose interval is widened for the number of looks taken (99.75% at the first look, 99.95% by the fifth). It supports M0, chance, if the whole interval lies below 51%, and M1, better than chance, if it lies above 50%;
- three Bayes factors comparing a hit rate of exactly 50% with a rate above 50%, each with a different prior for how
far above: the replication prior, built from Bem's Experiment 1 (828 of 1,560); a uniform prior; and a prior centred
near chance, beta 7 and 7, called the "BUJ" prior because, in the paper's words,
This knowledge-based prior was originally proposed by Bem, Utts and Johnson
. Each supports M0 if the evidence for chance exceeds 25 to 1, and M1 if it falls below 1 to 25.
The study stops for a model only when all four agree. Otherwise it goes on to the next look.
The four tests at the first registered look
All four say M0 at the first look: the rule stops. 18,876 hits in 37,836 erotic trials (49.89%) from 2,115 participants.
Recomputed here: 18,876 hits in 37,836 trials, 49.89%, from 2,115 participants
(the paper counts 2,220 people who took part; the rest gave no valid erotic trial before the stop, or started after it).
The paper printed 49.89% of 37,836 erotic trials from 2,115 participants, a 99.75% interval of
49.11% to 50.67% (here 49.107% to 50.671%). Its conclusion, which it says was
pre-written and approved during the consensus design process
, gives one Bayes factor without naming its prior:
Observing this percentage of successful guesses is 72 times more likely if the guesses are successful at random than if they have a better than chance success rate.Kekecs et al. (2023), Royal Society Open Science 10, 191375, section 5. Of this page's three Bayes factors only the BUJ one, 72.4, rounds to 72; the replication-prior 182.1 and the uniform-prior 212.3 are not in the paper's text and are this page's recomputation. So the prior Bem proposed with Utts and Johnson puts the evidence at 72 to 1 for chance.
Watch the curves while they run. The uniform-prior curve first passes 25 at trial 236 and stays above from trial 238; the replication and BUJ curves first pass it at trial 873 and stay above from trials 1,923 and 4,863. None of those crossings stopped anything. The rule looked only at 37,836, which is why a design like this can afford to be watched.
As run: 18,876 hits, 49.89%, 2,115 participants.
Two of the five excluded laboratory IDs were added to the analysis code about ten days into data collection, a
change the authors disclosed in their 2023 correction (it is not noted that two IDs to exclude were added 10 day after the data collection was started
) along with an auditor's
finding: The excluded data for the two added laboratory IDs were only 81 records widely dispersed throughout the study and could not affect the study conclusions.
Put the 19 erotic trials from those IDs back and the first look
holds 18,869 hits (49.87%) from 2,114 participants, with Bayes factors 192.8,
222.5 and 75.9. The verdict does not move. The same correction lists four further deviations
found by James E. Kennedy's audit of the project, and calls the first three potentially significant protocol deviations
: the
software's change history was kept on the project's own server instead of being pushed to GitLab in real time, so the research
auditors could not check during the study that the software was unaltered; the server access log was overwritten every few
days; and the IT auditor had ties to one of the collaborating laboratories. In the correction's words,
these protocol deviations are unlikely to have significant influence on the study conclusions, especially given that the study obtained a null result
.
The ten laboratories, as table 1 of the paper labels them
| Laboratory | Country | Participants | Erotic trials | Hit rate here |
|---|
Participant and trial counts match the paper's table 1 row for row; labels are the paper's, including "not disclosed". Hit rates are this page's.
III / The control on the control
Plant Bem's 53.1% into a copy of those trials.
A null result only counts if the test could have come out the other way. So take a copy of the replication's own 37,836 trials, choose some misses with a seeded generator (mulberry32, seed 20110131, a Fisher and Yates shuffle of the misses' positions), and turn each into a hit by setting its target side to the side the participant guessed. Nothing else changes: the order, the participants, the trial numbers and every guess stay as recorded. Then run the same, unmodified four-test rule on the copy.
Nothing planted yet. At the default 53.1%, the copy gets 1,215 edited curtains and 20,091 hits (53.10%), and the unmodified rule says M1 at the first look.
At the printed 53.1%, 1,215 of the copy's misses become hits. The interval moves to 52.32% to 53.88%, all three Bayes factors fall below 10⁻²⁹, and the rule stops for M1 at the first look. The real trials, through the same function, stop for M0. The control could have confirmed the claim. The other readings of the claimed size agree: 828 of 1,560 needs 1,206 edits and 826 of 1,560 needs 1,158; both stop for M1.
The knife edge is 51.0%, the smallest effect the replication said it cared about. That plant adds 420 hits (19,296 in all). The mixed model alone says M1 (interval 50.22% to 51.78%), but the uniform-prior Bayes factor is 0.041, just short of the 1/25 it needs, so the registered rule says inconclusive: not settled at the first look, and the design would have continued. That is neither a detection nor a miss, and the page will not call it either. Section VI answers what the full design would have done.
Planting only among Bem's stimulus seekers (the next section) puts 57.6% into the 15,997 trials of replication participants who meet his rule: 1,226 edits, an overall rate of 53.13%, and a spread between people that the mixed model now sees (its standard deviation rises from 0.052 to 0.158 on the logit scale). The rule says M1.
IV / The claimants' method, run on nothing
How often does Bem's test find the effect in coin flips?
Bem's inference was one t test over 100 sessions. Give it data known to contain nothing: 40 sessions of 12 fair coins and 60 of 18, which is exactly his structure, and the true null here, because the side was drawn after the guess. Then count how often the anchor's own function produces a t as large as his.
With seed 409: 60 of 10,000 null experiments (0.60%) reach t ≥ 2.51, against 0.68% from the t distribution; 498 (4.98%) reach one-tailed p < .05.
The procedure is calibrated. It does not manufacture the claim: by the t distribution a result at least as strong as Bem's turns up in about 1 in 146 experiments made of nothing, which is why one experiment alone could not settle it. The replication's own participants, cut to Bem's size (the first 2,100 in order of arrival, 21 blocks of 100), give 0 blocks with p < .05; their t values run from −2.16 to 1.23. These are the control's data, not data known to be empty.
The control's yardstick, turned on the claim
Run the replication's own Bayes factors on Bem's pooled erotic trials. With the uniform prior, the evidence for better than chance is 0.95 to 1; with the BUJ prior, 2.71 to 1 (with the replication's rounded 828: 1.21 and 3.45). Under the replication's own rule, Experiment 1 alone sits in the inconclusive band, nowhere near 25.
V / The second layer: the claim's strongest subgroup
Bem's largest effect was not 53.1%. Look for it where he said to.
Bem also reported that the effect lived in people high in stimulus seeking. His two-item scale was “I am easily bored” and “I often enjoy seeing movies I’ve seen before” (reverse scored), averaged
on five points. Participants scoring above the midpoint on the 5-point stimulus-seeking scale
hit 57.6% of the erotic trials, t(41) = 4.57, d = 0.71,
exact binomial p = .00008, and the correlation of the score with the hit rate was .18 (p = .035) (p. 410). Across his
experiments the stimulus seekers averaged d = 0.43.
The replication asked the same two questions. We included these questionnaires to match the original protocol by Bem as closely as possible
. However, data from these questionnaires are not used in hypothesis testing.
So the moderator sits in its released trials, untested. Scored as the replication's own script scores it (the first item
reversed), 2,111 participants average 2.71 (standard deviation 0.76), which is what the paper printed:
2.71 (0.76).
Cut: at least 3.0. Bem 42 sessions at 57.61% (t = 4.57); replication 894 participants at 49.94% (t = −0.15).
| Cut | Bem sessions | Bem rate | Bem t | Replication participants | Replication rate | Replication t |
|---|
The fork in the wording. The printed numbers (42 sessions, 57.6%) reproduce only with the rule "a score of 3.0 or more": 42 sessions at 57.61%, t = 4.57. Read "above the midpoint" strictly and the subgroup is 16 sessions at 55.56%, t = 2.04. The page shows both and infers nothing about why.
What the replication's participants do. Under Bem's operative rule, 894 of them qualify, with 7,988 hits in 15,997 erotic trials (49.93%) and a mean per person of 49.94%, t = −0.15. Under the strict wording, 509 qualify, at 49.35%, t = −1.22. The correlation of score with hit rate is 0.007, against Bem's 0.18. And the plant shows this is not a matter of power: put 57.6% into the qualifying participants' trials of the copy and the same t test on the same people returns 18.9.
We searched the Transparent Psi Project article, its 2023 correction, its stage 1 protocol and supplement, the 2025 AMP-TPP article, the Psi Encyclopedia entry (last updated 9 July 2026), the Wagenmakers et al. 2012 confirmatory replication report, James E. Kennedy’s 2023 audit of the project and its companion lessons document, and web searches on 2026-09-22 and did not find a published test of Bem’s two-item stimulus-seeking moderator in the Transparent Psi Project’s released trials. A 2022 item by the replication's authors in the Journal of Parapsychology (volume 86, pages 284 to 286) appeared in a search; only its first page could be seen, and it is left out of the list above for that reason.
VI / The second layer: where the replication was blind
For every true hit rate, what would the full design have said?
The real record ends at the first look. The design had four more. For a true hit rate p shared by everyone, this carries the exact distribution of hits through all five looks: at each look the thresholds on the hit count come from the same Bayes-factor and interval functions (at the first look, M0 for at most 18,993 hits and M1 for at least 19,297), and only the undecided probability goes on. No simulation. For the interval test this uses the version with no person-to-person spread, which is also what the real data look like: among the 2,087 participants who finished all 18 erotic trials the variance of hits is 4.55, against 4.5 for pure chance.
At a true rate of 50.70% the registered design stops for M1 with probability 86.84%.
| True rate | Printed by the replication | Exact, this page |
|---|---|---|
| 50.0% | false M1 lower than 0.0002 | false M1 0.0137%; M0 96.26% |
| 50.2% | false M0 in more than half of experiments | M0 70.40% |
| 50.7% | correct inference rate 88% | M1 86.84% |
| 51.0% | correct M1 greater than 0.95; false M0 lower than 0.001 | M1 99.87%; M0 0.0933% |
| 53.1% | not printed | M0 6.9 × 10⁻³⁰ |
The exact figures agree with the printed ones except at 50.7%, where the page gets 86.84% against the printed 88%. The replication ran 5,000 simulated experiments per scenario (10,000 at 50% and 51%), whose sampling error alone is about half a point, and fitted the mixed model where this page uses the interval without person-to-person spread. The page prints both numbers and tunes nothing. Choose "mixed model only" in section III and the curve and the readouts above change: that test alone gives a false M1 at 50% of 0.20%, which is why the design demanded agreement.
The blind zone is real, and the replication said so itself: extremely small effects might be unnoticed by our study
. It is also nowhere near the
claim. A true rate of 53.1% would have stopped for M0 with probability 6.9 × 10⁻³⁰.
VII / The verdict, dated
What the record says now.
As of 2026-09-22
DISSOLVED
For the claim quoted above: an above-chance hit rate on erotic trials in the Experiment 1 paradigm of Bem (2011). It says nothing about the other eight experiments of that article, and nothing about ESP in general.
The replication Bem helped design found 49.89% in 37,836 erotic trials; planted into a copy of those same trials, his 53.1% makes the unmodified four-test rule stop for it at the first look, so the control could have confirmed the claim, and did not.
Decided by independent replication, 12 years after the claim. The sources:
- Kekecs, Palfi, Szaszi and 27 coauthors (2023), Royal Society Open Science 10(2), 191375, doi:10.1098/rsos.191375, with its correction, Royal Society Open Science 10, 231080 (16 August 2023), doi:10.1098/rsos.231080.
- Walleczek, von Stillfried, Schmidt, Wittmann, Kirmse, Moll and Kekecs (2025), PLOS ONE 20(11), e0335330, doi:10.1371/journal.pone.0335330: three further studies with the same procedures, 26,483 participants and 420,472 critical trials. Study 1: 49.48%. Study 2: 49.65% in 127,000 trials, p = 0.013, below chance. Study 3: 50.07% in 217,800 trials, p = 0.496. In the authors' words,
none of the three replication studies was able to detect the precognitive effect (53.1%) reported by Bem
, andThe source of the one-time confirmed anomalous result in Study 2 remains to be identified.
- Michael Duggan, Transparent Psi Project, Psi Encyclopedia, Society for Psychical Research, last updated 9 July 2026. From the article body:
Including the earlier TPP study by Kekecs and colleagues, none of the four replications using TPP procedures found the above-chance precognitive effect reported by Bem
. Duggan is a coauthor of the claimants' 2016 meta-analysis.
Why DISSOLVED, and not ARTEFACT. The replication's authors went further than this page does: What we can conclude is that the original finding by Bem in this experiment is likely to be simply an artefact, and that this paradigm is unlikely to yield evidence of ESP if it does exist.
Critics have named objections that reach this experiment. James Alcock (Skeptical Inquirer, 2011) objected that its
procedure changed after 40 of the 100 sessions; Wagenmakers, Wetzels, Borsboom and van der Maas (Journal of Personality and
Social Psychology, 2011) argued that the article's analysis was partly exploratory and that one-sided p values may overstate
the evidence. Bem replied to both: he wrote that he divided the 100 sessions into two parts by design, to test several kinds of
nonerotic picture, and with Utts and Johnson he argued that a Bayesian analysis with a more reasonable prior gives strong
evidence for psi (that prior is the BUJ prior of section II). No source has shown that any of these objections produced the
53.1%, and "likely" is a judgement, not a finding. What the record shows is a result that more and better data did not reproduce. The
later below-chance result of AMP-TPP Study 2 is a different claim, in the other direction, whose source its authors
say is unidentified and which their own Study 3 did not reproduce; it is neither support for Bem nor a joke.
What would change it: a preregistered replication of this paradigm, with the same transparency, that the replication's own four-test rule (or one as strict) settles for M1; or a moderator present in Bem's Cornell sessions and absent from the replication (selected participants, an altered state, a stimulus set), shown prospectively to produce above-chance hit rates.
The claimants' side of the record
Asked in January 2018 whether he would comment on blog posts reanalysing his data, Bem wrote: I am happy to let replications settle the matter.
In the same message: Nor did I discard failed experiments or make decisions on the basis of the results obtained.
He also sat on the panel that designed the replication. His strongest later evidence is a meta-analysis with three
coauthors of 90 experiments on anticipating random future events (F1000Research, version 2, 2016): z = 6.40,
Hedges' g = 0.09, Bayes factor 5.1 × 10⁹. The replication notes that the design it replicated was the one that yielded the highest effect size in the 2016 meta-analysis
.
After the replication, its coauthor and meta-analysis coauthor Patrizio Tressoldi posted a preprint titled Further null evidence of extra-sensory-perception by using forced-choice experiments with unselected participants in an ordinary state of consciousness
(PsyArXiv, 25 August 2022), which reads the result as null for unselected participants and advises: Researchers of extra-sensory perception are advised to recruit skilled participants and request them to perform tasks in a modified state of consciousness as in ganzfeld field, dream, meditation, etc.
The Psi Encyclopedia reports a further objection: Dean Radin has argued that the kind of ‘hyper-objective’ approaches used in the TPP are counterproductive.
The replication's own discussion allows
that the positive literature might be the result of recognized methodological biases rather than ESP
, a sentence about the literature and not a finding about Experiment 1,
and in the next breath: the occurrence of ESP effects could depend on some unrecognized moderating variables that were not adequately controlled in this study, or ESP could be very rare or extremely small, and thus undetectable with this study design.
We found no published reply by Bem to the replication
(searched: the Psi Encyclopedia, the 2023 and 2025 papers, and the web, on 2026-09-22).
The correction also discloses: A small number (at least five out of 2115) of experimental sessions were restarted, thus, potentially leading to data from the same participant occurring in two research sessions instead of one.
In the released file one
participant ID carries 29 erotic trials within the first look; the registered analysis, and this page, treat it as one
participant.
The check
Your browser has not recomputed this page yet.
The data files have not been hashed yet.
| Quantity | As printed | Computed here | Tolerance | Result |
|---|---|---|---|---|
| Bem, erotic hit rate (session mean) | 53.1% | 53.139% | printed rounding | within |
| Bem, t(99) | 2.51 | 2.5133 | 0.005 | within |
| Bem, d | 0.25 | 0.25 | 0.005 | within |
| Bem, one-tailed p | .01 | 0.0068 | printed rounding | within |
| Bem, binomial z | 2.30 | 2.30 | 0.005 | within |
| Bem, erotic minus nonerotic t | 1.85 | 1.8563 | last digit: a gap, shown | last-digit gap |
| Bem, stimulus seekers, rule 3.0 or more | 57.6%, t = 4.57 | 57.61%, t = 4.57 | printed rounding | within |
| Replication, hit rate | 49.89% | 49.889% | printed rounding | within |
| Replication, participants | 2,115 | 2,115 | exact | within |
| Replication, 99.75% interval | 49.11% to 50.67% | 49.107% to 50.671% | rounds to printed | within |
| Replication, BUJ Bayes factor | 72 times | 72.4 | rounds to printed | within |
| Replication, guessed left | 49.08% | 49.08% | printed rounding | within |
| Replication, target left | 49.88% | 49.886% | last digit: a gap, shown | last-digit gap |
| Replication, sensation-seeking mean (s.d.) | 2.71 (0.76) | 2.7065 (0.7641) | printed rounding | within |
| Operating characteristic at 50.7% | 88% | 86.84% | 3 simulation standard errors, shown | within 3 simulation SEs |
The last column is evaluated, not typed: your browser applies each row's tolerance to the unrounded value it just computed and the number as printed, and would print "outside" where they part.
What the page could not do, and every free choice
- Bem's workbook holds one percentage per session, not trials, so his experiment cannot be replayed trial by trial; the extract converts percentages to counts with the printed design and refuses any row that does not divide exactly.
- The mixed model is this page's own Laplace fit, the approximation glmer uses by default, not R itself; its interval matches the printed one to the printed digits, and an independent fit by 25-point adaptive quadrature in the research directory moves its bounds by less than 0.001 percentage points.
- The operating characteristic assumes every participant shares one hit rate and uses the interval without person-to-person spread. With large person-to-person differences the replication's own simulations gave a 0.9 probability of correctly supporting M1 at 51%, against more than 0.95 without them.
- The plant works at the level of the released trials: it trusts that the replication’s software recorded guesses and targets as it says, and it cannot test stimuli, rooms or experimenters. It shows that the registered rule detects a 53.1% hit rate in these trials, not that an effect real under Bem’s conditions would have reached that size with the replication’s different erotic images and shared testing spaces.
- The plant chooses which misses to edit with a fixed seed. Another seed edits different curtains; the count, and so every Bayes factor, is the same, and the interval moves only in its last digits.
- Stimulus seeking is read, as the replication's script reads it, from each participant's first erotic row. Reading it from the first answered row of the file instead changes two participants' scores and makes the 3.0 group 893 rather than 894; the mean hit rate is 49.94% either way.
- Free choices, each a control on this page: the claimed size, the plant shape, which tests decide, the exclusion list, the stimulus-seeking cut, Bem's session subset, the true hit rate, and the null seed.
The verifier (research/thirty-seven-thousand-curtains/) recomputes every figure on this page in Node, against an independent recomputation in Python from the original OSF file and Bem's original workbook, and fails on any disagreement.