Ground truth · social psychology · the whole record

The Other Twenty-Two Rooms

Milgram did not run one obedience experiment. He ran 23 versions of it on 780 people, and the share who obeyed all the way to the last switch ran from 1 in 40 to 37 in 40. Here is every person in every room. Then rebuild, from its own counts, the 2014 paper that pooled them, down to a slip in one participant's code.

The number everyone learns comes from Stanley Milgram's first report, in 1963. Forty men from New Haven and the towns around it were told to shock a learner in an adjacent room for every wrong answer, starting at 15 volts and going “one step each time” up a board that ended at 450. The generator was a simulation, and the learner was “a 47-year-old accountant, trained for the role”. In Milgram's words, “Of the 40 subjects, 26 obeyed the orders of the experimenter to the end”. Twenty-six of forty is 65%.

Nobody had expected it. Before the results were known, fourteen Yale psychology seniors read a description of the setup and guessed how many of 100 people would go all the way: “The class mean was 1.2%.” Later, in his 1974 book, Milgram reported that psychiatrists he asked had predicted “about one subject in a thousand would administer the highest shock on the board.”

But that was one room. Over the following months Milgram kept changing the set: he moved the learner closer, sent the experimenter out, put in actors who refused, let the subject choose the voltage, moved the whole thing out of Yale to an office in Bridgeport. A 2014 paper that went back to his notes in the Yale archive counted them: “Milgram's famous experiment contained 23 small-sample conditions that elicited striking variations in obedient responding.” They were run “from August 1961 to May 1962”, and “Only 18 of these were reported in the monograph”, his 1974 book.

Every room

Each row below is one version of the experiment, numbered in the order it was run. Each dot is one person. A filled dot went to the 450-volt switch; an open one stopped somewhere before it. Tap a row to read what was different about that room, in the words of the paper that tabulated them all.

The 23 conditions, one dot per participant

    went to 450 Vstopped before it

    Some of what is in there. When the subject had to hold the learner's hand down on the shock plate, 12 of 40 went to the end. When two other “teachers” (actors) refused to go on, 4 of 40 did. When the experimenter left the voltage up to the subject, 1 of 40 did; when somebody else pressed the switch and the subject only read out the word pairs, 37 of 40 did. When the learner was the subject's own friend or relative, 3 of 20. The Bridgeport office, which textbooks like to cite, came out at 19 of 40, close to its Yale twin (condition 5, 26 of 40, or condition 6, the same procedure with different actors, 20 of 40).

    The same paper surveyed ten social psychology textbooks and found that “although the average text refers to 7.6 conditions, nine conditions go completely unmentioned.”

    Two cautions before reading the rows as laws. Each room held only 20 or 40 people, so every rate is rough: with 40 people, a room at 25 of 40 (62.5%) is consistent with a true rate anywhere from about 47% to 76% (a 95% Wilson interval, computed on this page). And the rooms were not a designed experiment that changes one thing at a time; the paper calls them “a patchwork of methodological elements rather than a systematic investigation”.

    Rebuild the paper that pooled them

    In 2014 Nick Haslam, Steve Loughnan and Gina Perry put all of these rooms into one analysis in PLoS ONE, “Meta-Milgram”. Their tables give, for each condition, how many people took part, how many went to 450 volts, and a set of yes-or-no codes for what was different about it: was the experimenter out of the room, did he give the orders, were there two experimenters who disagreed, did someone else press the switch, was the learner a friend. That is everything needed to redo their statistics, so the page does, in your browser.

    Their Table 4: obedience by each code, from the counts

    Twelve of the thirteen rows come back exactly. The thirteenth, “rights expression” (condition 8, where the learner says up front that he can leave when he wants), prints a rate of 0.41; the paper's own count for that room is 16 of 40, which is 0.40, and the chi-square it prints beside it, 0.23, is the one that 16 of 40 gives. A stray digit, and a harmless one.

    Their Table 5: the logistic regression, refitted

    Table 5 is the paper's main result: a logistic regression of going-to-450 on the codes, which finds eight properties of a room that predict obedience on their own. Refit it from the counts as the paper prints them and it comes close, but most rows miss in the second decimal: 2 of 16 match on B, SE and Wald.

    Now change one person. Take one participant who went to 450 volts in condition 22 (“Peer authority”) and give them the codes of condition 17 (“Teacher in charge”). The two rooms differ in exactly one code, whether the experimenter told the subject which shock to give, so this is the same as a single 0 typed where a 1 belonged. Refit, and 14 of 16 rows match the printed B, SE and Wald at the precision printed (the Wald values within rounding). The other two match on SE, Wald and p, and their printed B is the refit value with one zero dropped after the decimal point: the refit gives 0.006 for vulnerability where the table prints 0.06, and −0.059 for the quadratic proximity term where it prints −0.59. That is also why the table's Wald of 0.00 for vulnerability looked impossible next to a B of 0.06: with 0.006 it is right. And the table's p for low status, .614, is the Table 4 value for that row; the refit gives .301, which is the figure the paper's own text reports.

    How sure is this? The page did not go looking for one answer and stop. Of the 840 ways to move one participant from one room to another, and the 672 ways to give one participant one wrong code, exactly one reproduces Table 5, and it is this one. A wider search, over all 353,010 ways to move two participants, finds 31 more that work, and every one of them comes to the same thing once you net it out: one obedient person leaves condition 22 and arrives in condition 17, sometimes by way of a third room, sometimes alongside a swap between two rooms the regression cannot tell apart (conditions 5 and 6, and 18 and 19, carry identical codes).

    What it does not change: refit from the counts as published, every property the paper calls a significant predictor is still significant, and group pressure to obey is still only marginal. The paper's conclusions stand. What it cannot tell you: which participant, or whether the slip was in the data file, in the coding, or somewhere else on the way to the table. From outside, all anyone has is the printed numbers. No correction to the paper is listed at Crossref, PubMed or PLOS (checked 7 October 2026), and the authors have not been asked; this page is the first we know of to point it out, and that is all it claims.

    What the counts cannot tell you

    Every number above treats a room as a fixed procedure and a person as a 1 or a 0. The archive says neither was quite so.

    The script was not always the script. The experimenter had four set prods, from “Please continue, or Please go on” to “You have no other choice, you must go on.” Listening to the tapes of two conditions, the psychologist Stephen Gibson found that “the rhetorical strategies employed by the experimenter depart sometimes quite substantially from the official experimental script”, and that subjects could draw him into negotiation. Haslam and colleagues name this as a limit of their own analysis: the tapes show “that he often went beyond the standard ‘four prods’”.

    Not everyone believed it. Milgram's assistant Taketo Murata tabulated, in a study Milgram never published, how far subjects went against how much they believed the shocks were real. Reanalysing it in 2020, Gina Perry and colleagues report that “in 18 of 23 variations of the experiment, the mean levels of shock for those who fully believed that they were inflicting pain were lower than for subjects who did not fully believe”, and that subjects with a high level of belief “were 2.57 times more likely to be defiant than those who had a low level of belief.” If that holds, some of the obedience counted above was obedience to something the subject thought was staged.

    Where people stopped. A meta-analysis of eight of the conditions by Dominic J. Packer found that “In all studies, disobedience was most likely at 150 v, the point at which the shocked “learner” first requested to be released.” Jerry Burger used that point in 2009 to run a partial replication that went only as far as 150 volts: 70% of his base-condition subjects went on to the next item and had to be stopped, against 82.5% in Milgram's comparable condition, a difference that “fell short of statistical significance”.

    How this page is checked

    Every rate, interval, chi-square and regression on this page is computed in your browser by engine.mjs from conditions.mjs, which holds the paper's per-condition counts and codes, copied from its Tables 1 to 3 (the paper is CC BY). The check, verify-milgram-experiment.mjs, rebuilds both tables, refits the regression both ways, repeats the search over every one-person move and one-person miscoding, and finds each figure the prose quotes in the page's own sentences. The one figure it does not recompute is the two-person search (353,010 refits, about five minutes), which was run once while the page was written and whose result is kept with its research. Run node verify-milgram-experiment.mjs in an empty folder with Node 18 or later; it downloads the page's files and exits 0 if everything holds. --mutate breaks the engine on purpose and confirms the checks go red.

    It does not check that the sources say what this page says they say; the words relied on are kept beside each claim in the page's source record. Nor can it check the paper's counts against Milgram's own data sheets, which only the archive holds.

    Sources

    1. Haslam, N., Loughnan, S. & Perry, G. (2014). Meta-Milgram: An empirical synthesis of the obedience experiments. PLoS ONE 9(4): e93927. doi:10.1371/journal.pone.0093927 (open access, CC BY).
    2. Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology 67(4): 371–378. doi:10.1037/h0040525.
    3. Milgram, S. (1974). Obedience to Authority: An Experimental View. New York: Harper & Row. Chapter 3, “Expected Behavior”, p. 31.
    4. Gibson, S. (2013). Milgram's obedience experiments: A rhetorical analysis. British Journal of Social Psychology 52: 290–309. doi:10.1111/j.2044-8309.2011.02070.x.
    5. Perry, G., Brannigan, A., Wanner, R. A. & Stam, H. (2020). Credibility and incredulity in Milgram's obedience experiments: A reanalysis of an unpublished test. Social Psychology Quarterly 83(1): 88–106. doi:10.1177/0190272519861952.
    6. Packer, D. J. (2008). Identifying systematic disobedience in Milgram's obedience experiments: A meta-analytic review. Perspectives on Psychological Science 3(4): 301–304. doi:10.1111/j.1745-6924.2008.00080.x.
    7. Burger, J. M. (2009). Replicating Milgram: Would people still obey today? American Psychologist 64(1): 1–11. doi:10.1037/a0010932.