The Test for Intellectual Blooming
In 1966 two researchers reported that children whose teachers had been told they would bloom, although they had been picked at random, gained more IQ than their classmates. Recomputed from the children’s own scores, the first and second graders named to their teachers gained 11.0 more total IQ points than the rest; this page shows that first. Then it runs the critics’ test, which drops every score the IQ test was never normed for: the advantage stays at 9.9 points. Planted at the claimed size into re-drawn lists of real children, that test keeps the effect in 94.1% of copies for total IQ but only 31.0% for reasoning IQ, so the reasoning half of the critique is inconclusive, and whether expectations raise IQ stays open.
The first and second graders, before and after
Oak School, May 1964 to May 1965. The first and second graders, IQ before and after the year their teachers were handed a list of bloomers.
Total IQ, all scores: the 19 named children gained 11.0 points more than the 95 others (two-tailed p = 0.004).
Each point is one child: pretest in May 1964 across, basic posttest in May 1965 up. The five children Snow picked out in 1995 by their reasoning scores are ringed, with their listing IDs. A point above the diagonal gained.
From the children’s own scores, the first and second graders named to their teachers gained 11.0 more total IQ points than the others; dropping every score outside the test’s norms leaves 9.9 (two-tailed p = 0.002), and planted at the claimed size into re-drawn lists of real children that deletion keeps the effect in 94.1% of copies; for reasoning IQ it keeps it in only 31.0%, so the reasoning half of the critique is inconclusive, and whether expectations raise IQ remains open.
Everything below is computed in your browser from two small data files, the listing of the children’s scores that Rosenthal and Jacobson supplied to their critics, transcribed by this page, and nineteen rows of later experiments, plus a third file holding the printed numbers they are compared against. Numbers taken from a publication are marked as printed and cited where they appear; every test, p-value, power figure and re-drawn list is the page’s own and says so. The check at the bottom recomputes the page against itself while you read it.
I · the claim, at full strength
Eight months later, the bloomers had gained more
In May 1964 every child at an elementary school in South San Francisco, called Oak School in the reports, sat a group intelligence test. It was Flanagan’s Tests of General Ability, TOGA, which gives a verbal IQ, a reasoning IQ and a total. The teachers were told something else: that it was a test for intellectual blooming, and that its scores named the children who would spurt ahead in the coming year. In the autumn each of the eighteen teachers received a list of such children. The names had been chosen, in the authors’ words, “by means of a table of random numbers”. The school retested the children in January 1965, May 1965 and May 1966. Robert Rosenthal and Lenore Jacobson reported the result in August 1966:
Within each of 18 classrooms, an average of 20% of the children were reported to classroom teachers as showing unusual potential for intellectual gains. Eight months later these “unusual” children (who had actually been selected at random) showed significantly greater gains in IQ than did the remaining children in the control group. These effects of teachers' expectancies operated primarily among the younger children.
Robert Rosenthal and Lenore Jacobson, “Teachers’ Expectancies: Determinants of Pupils’ IQ Gains”, Psychological Reports 19 (1966), 115 to 118, doi:10.2466/pr0.1966.19.1.115; the summary as deposited with the publisher’s record
Their 1968 book, Pygmalion in the Classroom, gave fuller tables. Then they did something that let this page exist: they supplied the children’s scores, card by card, to two researchers at Stanford, Janet Elashoff and Richard Snow, who printed all of it as an appendix to their December 1970 report. That listing holds 382 children, the ones present for at least one posttest: 126 in grades 1 and 2, 131 in grades 3 and 4, 125 in grades 5 and 6, each with a grade, a track (fast, medium or slow), a code for named or not, and total, verbal and reasoning IQ at four testings. This page transcribed it twice, from the scan, and reconciled the two passes cell by cell (the check, below). Here is the claim recomputed from the children’s own scores, gain from May 1964 to May 1965, beside the numbers the authors printed.
| grade | not named, N | gain | named, N | gain | advantage | t | as printed | N as printed |
|---|---|---|---|---|---|---|---|---|
| 1 | 48 | 12.0 (16.6) | 7 | 27.4 (12.5) | +15.4 | 2.98 | +15.4, t 2.97 | 48 / 7 |
| 2 | 47 | 7.0 (10.0) | 12 | 16.5 (18.6) | +9.5 | 2.29 | +9.5, t 2.28 | 47 / 12 |
| 3 | 40 | 5.0 (11.9) | 14 | 5.0 (9.3) | 0.0 | 0.0 | 40 / 14 | |
| 4 | 49 | 2.2 (13.4) | 12 | 5.6 (11.0) | +3.4 | +3.4 | 49 / 12 | |
| 5 | 26 | 17.5 (13.1) | 9 | 17.4 (17.8) | 0.0 | −0.1 | 26 / 9 | |
| 6 | 45 | 10.7 (10.0) | 11 | 10.0 (6.5) | −0.7 | −0.7 | 45 / 11 | |
| all | 255 | 8.42 (13.5) | 65 | 12.22 (15.0) | +3.80 | 2.13 | +3.80, t 2.15 | 255 / 65 |
Every mean and every standard deviation the authors printed comes back from the listing to the last digit they printed. (The 1966 sigmas turn out to be population standard deviations.) Across the school the 65 named children gained 12.22 points and the 255 others 8.42, an advantage of 3.80; the book prints 12.22, 8.42 and 3.80. The authors tested it against a mean square within classrooms that they print as 164.24. The listing gives exactly 164.24, but only when the degrees of freedom are counted the way their analysis of variance counted them: 320 children minus all 36 cells of a six-grade, three-track, two-group design, including the two empty cells of the fifth-grade class in which part of the test was not re-administered, which leaves 284. On that mean square the school-wide t is 2.13; the paper prints 2.15. The page cannot recover the last 0.02, and it does not matter to the claim as printed: the one-tailed p is 0.017, which rounds to the printed .02.
The advantage lives in the youngest classes. In grade 1 it is 15.4 points (t = 2.98; printed 2.97), in grade 2 9.5 (t = 2.29; printed 2.28), and in grades 3 to 6 between −0.7 and 3.4. Taken together, the first and second graders named to their teachers gained 20.53 points against 9.53, an advantage of 11.00. The best first-grade classroom, the slow track, shows the named children 24.75 points ahead of classmates who gained 16.25 (printed: 24.8 over 16.2); the best second-grade classroom, the fast track, 18.24 ahead of 4.26 (printed 18.2 over 4.3).
Three printed statistics are subtler. The book’s F of 6.35 for the whole school comes back as 6.35 from an unweighted-means analysis of grade by treatment with the same mean square, the approximation Elashoff and Snow say the authors used for unequal cells; it is not the same test as the t above. The book’s correlation between grade and advantage, r = −.86, comes back as −0.86. The 1966 rank correlation, rho = −.94, does not: in the listing grade 3’s advantage (−0.03) sits just below grade 5’s (−0.02), the rounded table prints them the other way round (0.0 and −0.1), and the unrounded rank correlation is −0.83. The 1966 correlation of the named and unnamed children’s gains across the seventeen classrooms, .57, returns as 0.57.
Ten, twenty, thirty points
For the first two grades the authors also counted children who gained at least 10, 20 or 30 points (1966 Table 2). From the listing: at least 10, 15 of the 19 named children (78.9%) and 46 of the 95 others (48.4%); at least 20, 9 (47.4%) and 18 (18.9%); at least 30, 4 (21.1%) and 5 (5.3%). The three chi-squares with Yates’ correction come back as 4.77, 5.59 and 3.47, against a printed 4.75, 5.59 and 3.47. One printed share does not: the paper gives 49% of the unnamed children gaining 10 points, and the listing gives 48.4%. The printed chi-square fits the listing’s count, so either the paper rounded 48.4% up or one child’s count differs; the listing alone cannot say which.
Verbal and reasoning IQ
Total IQ is built from two subtests, and the book prints both. Across the school most of the total’s advantage comes from the reasoning items (7.13 points against 2.06 for verbal); in the first two grades the two subtests move about equally (12.66 and 10.03).
| subtest and grades | not named, N | gain | named, N | gain | advantage | t, p | as printed |
|---|---|---|---|---|---|---|---|
| verbal, grades 1 and 2 | 95 | +4.49 | 19 | +14.53 | +10.03 | 2.24, 0.013 | +4.5 / +14.5 / +10.0, p .02 |
| verbal, grades 3 to 6 | 174 | +9.59 | 49 | +8.04 | −1.55 | −0.54, 0.705 | +9.6 / +8.0 / −1.6 |
| verbal, all | 269 | +7.79 | 68 | +9.85 | +2.06 | 0.85, 0.197 | +7.79 / +9.85 / +2.06 |
| reasoning, grades 1 and 2 | 95 | +26.97 | 19 | +39.63 | +12.66 | 1.95, 0.026 | +27.0 / +39.6 / +12.7, p .03 |
| reasoning, grades 3 to 6 | 160 | +9.06 | 46 | +15.93 | +6.88 | 1.59, 0.056 | +9.1 / +15.9 / +6.9, p .06 |
| reasoning, all | 255 | +15.73 | 65 | +22.86 | +7.13 | 1.99, 0.024 | +15.73 / +22.86 / +7.13, p .005 |
The mean squares within classrooms come back as 316.40 for verbal and 666.58 for reasoning, where the book prints 316.40 and 666.58. The verbal row needed one cell of the transcription settled by this table, and the check says which. For reasoning across the school the book prints a one-tailed p of .005; the t on the printed mean square gives 0.024, and only the unweighted-means F (6.98) gives a p that small, so that row appears to rest on the F. For the first and second graders the claimants’ test, on the whole school’s mean square as the book computes it, gives t = 2.24 for verbal IQ (one-tailed p = 0.013; the book, p. 77: t = 2.24, p < .02) and t = 1.95 for reasoning IQ (p = 0.026; printed .03).
The claimants’ test, as the page runs it
Gain is the chosen posttest minus the May 1964 pretest; the mean square counts every cell of the design, as the book’s does. The listing holds no May 1966 score for any sixth grader, so that combination is refused rather than computed from nothing.
II · the deciding control
The scores the test was never built to give
Elashoff and Snow began with the instrument. TOGA turns a child’s raw score into an IQ through tables in its manual, and those tables, they found, stop at 60 and 160:
However, the tables showing IQ scores for each raw score and age are not extrapolated beyond IQs of 60 and 160.
Elashoff and Snow, 1970 report, p. 38
The listing is full of scores outside that range. Of its 3,962 IQ scores, 150 lie outside 60 to 160, among them 11 scores of exactly zero and 15 above 200. These are scores outside the range the test was normed for, which is not the same as wrong; but a gain computed from them rests on arithmetic the manual does not support. Elashoff and Snow offered two remedies:
One procedure is to truncate the data by excluding as too poorly measured any IQ scores outside this range. Another possibility is renorming the data by replacing all scores less than 60 by 60 and all scores higher than 160 by 160.
Elashoff and Snow, 1970 report, p. 61
They ran both, for each pair of grades that took one TOGA form (grades 1 and 2, 3 and 4, 5 and 6) separately, on four criteria: the pretest difference, the posttest difference, the gain, and the posttest adjusted for the pretest. The panel below is that analysis, and every cell of their Tables 20 to 22 comes back from the listing: 108 advantages, 103 of them within 0.10 points of print and the other 5 within 0.14, and the significance mark (two-tailed p < .05) the same in 108 of 108. Their conclusion:
Our reanalysis reveals no treatment effect or “expectancy advantage” in grades 3 through 6. The first and second graders may or may not exhibit some expectancy effect; these experimental and control groups differ greatly on the pretest and a statistical analysis of such data cannot provide clear conclusions. There is enough suggestion of an expectancy effect in grades 1 and 2 to warrant further research, but the RJ experiment certainly does not demonstrate the existence of an expectancy effect or indicate what its size may be.
Elashoff and Snow, Pygmalion Reconsidered (1971), p. 44
The deciding control: delete or renorm the scores outside the norms
Covariance across all six grades is refused, because the three grade pairs took three different TOGA forms and Elashoff and Snow warn: “Covariance analysis or gain score analysis using all grades is unwise because of the dissimilarity in pre-posttest relationships across grades.” (Elashoff and Snow, 1970 report, p. 80). The same sentence cautions against a gain analysis across all grades, which is the claimants’ own analysis; the page runs that one because it is the claim.
Every cell of Elashoff and Snow’s Tables 20 to 22, recomputed (top) beside print (below)
| IQ, grades, scores | pretest | posttest | gain | adjusted |
|---|---|---|---|---|
| total, grades 1 and 2, all | 4.954.9 | 15.95*15.9* | 11.00*11.0* | 12.83*12.8* |
| total, grades 1 and 2, renormed | 4.474.5 | 13.74*13.7* | 9.26*9.2* | 10.88*10.8* |
| total, grades 1 and 2, truncated | 0.670.7 | 10.60*10.6* | 9.93*9.9* | 10.16*10.1* |
| total, grades 3 and 4, all | 0.480.5 | 2.302.3 | 1.821.8 | 1.862.0 |
| total, grades 3 and 4, renormed | 0.480.5 | 2.102.1 | 1.621.6 | 1.661.6 |
| total, grades 3 and 4, truncated | −1.83−1.9 | 0.140.1 | 1.972.0 | 1.721.7 |
| total, grades 5 and 6, all | 4.304.3 | 4.484.5 | 0.180.2 | −0.01−0.1 |
| total, grades 5 and 6, renormed | 4.244.3 | 4.384.4 | 0.140.1 | −0.030.1 |
| total, grades 5 and 6, truncated | 3.643.6 | 2.342.3 | −1.30−1.3 | −1.39−1.4 |
| verbal, grades 1 and 2, all | 0.460.4 | 10.49*10.5* | 10.03*10.1* | 10.15*10.2* |
| verbal, grades 1 and 2, renormed | 0.540.5 | 8.93*9.0* | 8.39*8.5* | 8.58*8.7* |
| verbal, grades 1 and 2, truncated | −1.43−1.4 | 6.896.9 | 8.33*8.3* | 7.83*7.8* |
| verbal, grades 3 and 4, all | 4.064.0 | −0.62−0.6 | −4.68−4.6 | −4.86−4.8 |
| verbal, grades 3 and 4, renormed | 3.253.2 | −3.63−3.6 | −6.88*−6.8* | −6.44*−6.4* |
| verbal, grades 3 and 4, truncated | −1.75−1.7 | −7.39−7.3 | −5.64−5.6 | −5.93−5.9 |
| verbal, grades 5 and 6, all | 0.690.7 | 2.782.7 | 2.102.0 | 2.052.0 |
| verbal, grades 5 and 6, renormed | 0.720.7 | 1.051.0 | 0.320.3 | 0.380.4 |
| verbal, grades 5 and 6, truncated | 3.033.0 | 1.601.6 | −1.43−1.4 | −1.05−1.0 |
| reasoning, grades 1 and 2, all | 13.2213.2 | 25.88*25.8* | 12.6612.6 | 21.10*21.0* |
| reasoning, grades 1 and 2, renormed | 8.408.4 | 18.59*18.6* | 10.1910.2 | 13.64*13.7* |
| reasoning, grades 1 and 2, truncated | 0.300.3 | 5.976.0 | 5.675.7 | 5.805.8 |
| reasoning, grades 3 and 4, all | −3.02−3.0 | 5.715.7 | 8.748.7 | 8.388.3 |
| reasoning, grades 3 and 4, renormed | −3.01−3.0 | 6.306.3 | 9.31*9.3* | 8.52*8.5* |
| reasoning, grades 3 and 4, truncated | −3.44−3.4 | 6.896.9 | 10.33*10.3* | 9.05*9.0* |
| reasoning, grades 5 and 6, all | 4.094.0 | 8.878.9 | 4.784.8 | 5.425.4 |
| reasoning, grades 5 and 6, renormed | 4.094.1 | 3.823.9 | −0.27−0.2 | 0.460.5 |
| reasoning, grades 5 and 6, truncated | 3.193.2 | −1.60−1.6 | −4.79−4.8 | −4.03−4.0 |
The largest gap from print in each column: pretest 0.09, posttest 0.09, gain 0.11, adjusted 0.14. In the pretest, posttest and gain columns, 72 of the 81 printed values are exactly the difference of two group means each first rounded to one decimal; the page keeps the unrounded difference, so a printed 9.2 can come back as 9.26. More than 0.10 from print: 5 cells, 4 of them in the adjusted column (total IQ grades 3 and 4 with all scores, 1.86 against 2.0; total IQ grades 5 and 6 renormed, −0.03 against 0.1; verbal IQ grades 1 and 2 renormed, 8.58 against 8.7; and reasoning IQ grades 1 and 2 with all scores, 21.10 against 21.0), and the verbal IQ grades 1 and 2 renormed gain, 8.39 against 8.5. Rounding does not explain these: a different covariance specification might, or verbal scores slightly different from the listing’s (Elashoff and Snow print the verbal gain advantage in grades 1 and 2 as 10.1 where Rosenthal and Jacobson print 10.0 for the same comparison), and the page cannot tell which. None of them changes a significance mark. The adjusted column is an analysis of covariance with one slope within the two groups.
For total IQ in the first two grades the deletion does not remove the advantage. With every child dropped whose pretest or posttest lies outside 60 to 160, the named children still gained 9.9 points more (two-tailed p = 0.002; printed 9.9), and adjusted for the pretest 10.2 (p = 0.0006). Of the nine ways the panel can score the first and second graders’ total IQ after the pretest (three treatments of the extreme scores, three criteria), 9 show a significant advantage for the named children.
Reasoning IQ is another matter. There the named first and second graders started 13.2 points ahead on the pretest, gained 12.7 more (two-tailed p = 0.129 by the critics’ pooled test; one-tailed p = 0.026 by the claimants’ test on the whole school’s mean square, which the book prints as .03), and after truncation, which keeps 12 named and 62 unnamed children, the advantage is 5.7 with p = 0.365. Of the nine reasoning analyses, 4 are significant, none of them truncated. Snow put it in one line in 1995, writing about his Figure 1, the reasoning-IQ scatter of the first and second graders:
The expectancy effect disappears when extreme scores are omitted.
Snow, 1995, p. 170
He wrote that the named children’s higher line “appears to result solely from five children whose respective pretest-posttest scores were 17-110, 18-122, 133-202, 111-208, and 113-211.” Those are, in the listing, IDs 10 and 1 in the medium first-grade class, and IDs 106, 139 and 119 in the second grade, ringed on the scatter above. He also wrote: “About 35% of the scores fall outside the norm range.” In the listing, 40 of the 114 first and second graders (35.1%) have a reasoning pretest or posttest outside 60 to 160, while 43 of their 228 reasoning scores (18.9%) do. Snow’s 35% matches the share of children; this is the page’s reading of his sentence, not his. For total IQ the same share of children is 5.3%, which is why the two halves of the critique part company.
Snow went on to suggest that teachers who gave the tests may have coached some children (“consider again the five outlying asterisks in Figure 1!”). That is a suggestion, not a finding, and the authors had tested it in 1966: “three of the classes were retested by a school administrator not attached to the particular school. She did not know which children were in the experimental condition.” Her results did not differ significantly from the teachers’, and “there was a tendency for the results of her retesting to yield even larger effects of teachers’ expectancies.”
The claimants’ reply, and its evidence
Rosenthal and Donald Rubin answered in the same book: “Despite the varied procedures employed in ES, the expectancy effects found in RJ remain undiminished.” Their evidence was their Table 31: for total IQ, by grade group, the claimants’ own gain score beside Elashoff and Snow’s eight other scores (posttest, gain and adjusted, each with all, renormed and truncated scores). For grades 1 and 2 it prints the claimants’ 11.0 (95% interval 4.7 to 17.3), the eight others from 9.2 to 15.9, and 9 of the nine significant; for grades 3 and 4, 0; for grades 5 and 6, 0. From the page’s own cells the same table comes back: 11.0 (4.7 to 17.3), others from 9.3 to 15.9, and 9, 0 and 0 significant. They concluded: “The results of the varied ES analyses are absolutely consistent with the results of the RJ analyses and indicate a significant effect of teacher expectations.” Elashoff and Snow’s rebuttal: “It is heartening that RR now admit no effects beyond the first two grades, since the text of RJ’s report implied generally significant results.” In 1995 Rosenthal wrote: “Among their reanalyses of the original data, they tried eight variations, including some that were statistically biased. Unfortunately for their position, every one of their reanalyses supported the original conclusions of the Pygmalion study.” The panel is where a reader can see which half of that exchange each number supports: for total IQ in grades 1 and 2, every score treatment keeps the advantage; for reasoning IQ, truncation does not.
The other experiments
The claim has a second half: not whether Oak School’s numbers show an advantage, but whether the effect exists. Other experiments tried to induce expectations in teachers and measure IQ, and in 1984 Stephen Raudenbush synthesised eighteen of them, Oak School among them:
It was hypothesized that the better teachers know their pupils at the time of expectancy induction, the smaller the treatment effect would be. The data strongly supported this hypothesis.
Raudenbush, 1984, abstract
The machine-readable rows are Raudenbush and Bryk’s 1985 coding, as the metadat package ships it. The page reproduces the package’s printed random-effects fit of all nineteen rows: 0.0837 (printed 0.0837), standard error 0.0516, z = 1.62, p = 0.105, between-study variance 0.0188, I² 41.85% (printed 41.86%), and for the meta-regression on weeks of contact capped at three, intercept 0.407, slope −0.157, QM 19.259 (printed 19.258). Without the Oak School row, which should not vote in its own control, the other 18 rows (seventeen experiments; Pellegrini and Hicks appear twice, by tester) pool to 0.057 (z = 1.22). Split where Raudenbush split them, at two weeks of teacher contact before the induction, the 10 low-contact rows pool to 0.274 (z = 2.81) and the 8 high-contact ones to −0.063 (z = −1.26). The 1984 paper’s own coding gives 0.23 and −0.06 (and 0.11 overall); the 1985 rows differ from it in places, most visibly for Pellegrini and Hicks, which Snow said the 1984 synthesis entered as 0.52. Applying Snow’s objection, only the tester-blind Pellegrini and Hicks row, the low-contact estimate is 0.198 (z = 2.53).
The other experiments, as the page pools them
Each row is one experiment: its first author and year, then its weeks of teacher contact before the induction, and “blind” (“b” on a narrow screen) where the tester did not know which children had been named.
III · the control on the control
Could deleting the extreme scores have kept a real effect?
A deletion that removes an effect proves nothing unless it could have kept one. Rosenthal and Rubin argued that this one would tend to lose a real effect: “On the other hand, when these procedures are applied to posttest scores they are biased and tend to diminish any real differences between the experimental and control groups.” The listing lets the page test the argument with the real children. It takes the 320 children who had a pretest and a basic posttest, drops the named ones, and in each first- and second-grade classroom draws at random as many of the remaining children as that classroom really had named: pseudo-bloomers from children the claim itself says received no expectancy. It adds the claimed advantage to each pseudo-bloomer’s May 1965 score, the printed 15.4 points in grade 1 and 9.5 in grade 2 (Table 7-1), unrounded and unclipped. Then it runs Elashoff and Snow’s truncated test, unmodified, and asks whether it finds a positive advantage at two-tailed p < .05, their own mark. It does this 2,000 times, with seeds 1 to 2,000. The page counts the control able to confirm the claim when it keeps the planted effect in at least 80% of the copies, the power a study is conventionally designed to have; below that it calls the control inconclusive at that size.
Plant the claim into re-drawn lists of real children
For total IQ the answer is yes. At the claimed size the truncated test finds the planted advantage in 94.1% of the copies, where the original all-scores test finds it in 86.5%; with nothing planted it finds a positive advantage in 0.5%. Truncation deleted about 1.0 of the 19 pseudo-bloomers per copy, and on average 0.00 of them because the plant pushed a posttest over 160. So the control COULD have confirmed the total-IQ claim, and on the real list it kept the advantage: 9.9 points.
For reasoning IQ the answer is no. At the claimed 12.7 points the truncated test finds the planted advantage in only 31.0% of the copies, and the all-scores test in 18.0%; truncation deletes about 6.2 of the 19 pseudo-bloomers per copy, 0.31 of them because of the plant. The control could NOT have confirmed the reasoning-IQ claim: a real effect of the claimed size would have “disappeared” from it in 69.0% of copies. Snow’s sentence is therefore inconclusive at the size claimed, and the page does not read it as showing that the reasoning result was produced by the extreme scores. The measurement is more to blame than the deletion: reasoning gains in the first two grades are so spread out that with every score kept the same test finds a real effect of that size in only 18.0% of copies.
The other experiments get the same treatment. Planting the dataset’s own coding of Oak School, d = 0.30, into every row but Oak’s moves the pooled estimate from 0.057 to 0.357 (z = 7.58); planted into only the 8 high-contact experiments it moves their estimate from −0.063 to 0.237 (z = 4.74). The high-contact experiments could have seen an Oak-sized effect and did not; the smallest planted effect they would have shown at z ≥ 1.96 is d = 0.16. (The page’s own d for Oak School, the advantage divided by the square root of the mean square, is 0.296; divided by the standard deviation of all 382 pretests, 18.48, as Rosenthal and Rubin did, it is 0.21; Rubie-Davies and Hattie give 0.35, without saying how it was computed.)
Plant an effect into the other experiments
IV · the claimants’ method on nothing
The claimants’ test, run on children nobody named
The same re-drawn lists answer a second question: how often does the authors’ own test, the function that reproduced their table above, find their result among children nobody named? The page draws pseudo-bloomers in all seventeen classrooms from the 255 unnamed children, 10,000 times, plants nothing, and runs the test.
The claimants’ test on untreated children
A school-wide t of 2.15 or more turns up in 4.5% of the lists (450 of 10,000), where the t table promises 1.6% on 219 degrees of freedom: about 2.8 times too often. At least one grade reaches the printed first-grade level, one-tailed p ≤ .002, in 5.9%, where six independent, well-calibrated tests would do so in 1.2%. Two features of the design do this, both visible in the draws. The classrooms had very different shares of named children and very different average gains, so a school-wide comparison of all named with all unnamed children carries an offset even when nothing is planted: 0.99 points in expectation (0.94 across these lists), a mean t of 0.52. And first graders’ gains vary far more than the pooled mean square assumes, so first-grade t values spread with a standard deviation of 1.35 instead of 1. So at the level of the school the test overstated the evidence: the listing’s own t of 2.13 is reached or passed in 4.6% of these lists, and the re-drawn lists below give p = 0.029, where the t table gave 0.017. It did not produce the first and second graders’ advantage: an advantage of 10.99 points or more, the size Table 7-1 implies for those grades, appeared in 7 of 10,000 lists. This is the page’s computation on the claimants’ untreated children, not data from an independent null.
V · the second layer
Draw the bloomers again
The names were drawn from a table of random numbers, so the most direct test of the list is the drawing itself. The page pools every child with the measure at both testings, re-draws the named children within each classroom with the real counts, 10,000 times, and asks how often a list drawn by chance does as well as the real one. That needs no assumption about the shape of scores running from 0 to 300; it does assume the drawing was within classrooms and nothing finer. Elashoff and Snow pointed out that the authors described their randomisation only as done “within blocks of classrooms”, perhaps balanced by sex too, which this page does not transcribe; a finer drawing would change the reference set.
Re-draw the list
| IQ, scores | school-wide advantage | p | grades 1 and 2 advantage | p |
|---|---|---|---|---|
| total, all | 3.80 | 0.029 | 11.00 | 0.002 |
| total, renormed | 3.23 | 0.044 | 9.26 | 0.001 |
| total, truncated | 3.17 | 0.042 | 9.93 | 0.0005 |
| reasoning, all | 7.13 | 0.015 | 12.66 | 0.035 |
| reasoning, renormed | 5.85 | 0.012 | 10.19 | 0.013 |
| reasoning, truncated | 3.69 | 0.147 | 5.67 | 0.136 |
For the school as a whole, the real list of total-IQ gains beats re-drawn lists with p = 0.029, larger than the 0.017 the t test gave, because re-drawn lists carry the same kind of offset described above (0.51 points in expectation here, 0.48 across these lists). For the first and second graders the real list stands far out: p = 0.002 with all scores, 0.001 renormed, 0.0005 truncated. For reasoning IQ in those grades, p = 0.035 with all scores and 0.136 truncated.
How much could Oak School see?
The last instrument turns the question around. Plant a uniform advantage from 0 to 20 points into pseudo-bloomers in grades 1 and 2, 2,000 lists at each size, and run both the claimants’ test and the truncated control. At the size Table 7-1 implies, 10.99 points, they find it 90.6% and 91.8% of the time. The truncated control first reaches 80% at 10 points. The other experiments suggest a much smaller effect where teachers had known their pupils for two weeks or less: the low-contact estimate above, converted to points at the square root of the mean square (the convention that maps the dataset’s Oak d of 0.30 to about 3.8 points), is 3.5 points, and at that size Oak School would have seen it only 18.9% of the time with the claimants’ test and 10.5% with the truncated one. (On the curve itself, the low-contact mark follows the other experiments’ panel as it is set when the curve is computed; the numbers in this paragraph are for its default two-week cut.)
The power curve of Oak School
We searched the text of Elashoff and Snow’s 1970 report and 1971 book (with Rosenthal and Rubin’s reply and the authors’ rebuttal), Snow 1995 (with the first page of Rosenthal 1995), Spitz 1999, Jussim and Harber 2005, Rubie-Davies and Hattie 2024, Raudenbush 1984, the metadat documentation, the Artificial Wasteland corpus, and web search results for “Pygmalion in the Classroom Oak School randomization test OR permutation test reanalysis statistical power” and “Pygmalion Rosenthal Jacobson reanalysis raw data TOGA norm range 60 160 extreme scores interactive” on 2026-09-23 and did not find a within-classroom re-randomisation test of the Oak School listing, or a calculation of whether the norm-range deletion could have kept an effect of the claimed size.
VI · the verdict
Open: a firm number, an unsettled question
OPEN
As of 23 September 2026 (2026-09-23), on the question of whether expectations induced in teachers raise young children’s IQ scores. Not decided.
The category and the scope come from the sources below, and the date is when this page checked them; the review that calls the IQ question unresolved also calls the large and dramatic version disconfirmed. A small effect on young children, induced before their teachers know them well, is neither established by the record nor excluded by it.
A 2005 review of the whole teacher-expectation literature put it this way:
It therefore appears that whether teacher expectations have much influence on student intelligence remains controversial and unresolved.
The hypothesis that teacher expectations have large and dramatic effects on IQ has been disconfirmed.
Jussim and Harber, 2005, p. 137
If one believes the critics, the IQ effect is zero. If one believes the advocates, it is very small (frequently 0, never consistently much higher than an r of .2).
Jussim and Harber, 2005, p. 152
And a 2024 review still reports the Oak School result as a finding:
By the end of the academic year, the ‘bloomers’ did achieve higher IQ scores than those for whom high expectations had not been induced (d = 0.35).
Although this study had many critics (e.g. Thorndike 1968; Spitz 1999), it was the catalyst for a new field in educational and social psychology.
Christine M. Rubie-Davies and John A. Hattie, “The powerful impact of teacher expectations: a narrative review”, Journal of the Royal Society of New Zealand 55(2) (2024), 343 to 371
What this page adds, labelled as its own: from the children’s scores, the first and second graders’ total-IQ advantage survives the deletion of every score outside the norms, the deletion had the power to keep it, and a re-drawn list almost never matches it. That is a statement about Oak School’s numbers. It is not a statement that expectations raise IQ: the reasoning-IQ half of the critique is inconclusive in both directions, the school-wide significance was inflated by the test, the pretest of the young named children already ran ahead, the teachers gave the tests, and the seventeen other experiments support at most a small effect, only where teachers had known their pupils for two weeks or less.
Sources for the verdict
- Lee Jussim and Kent D. Harber (2005). Teacher Expectations and Self-Fulfilling Prophecies: Knowns and Unknowns, Resolved and Unresolved Controversies. Personality and Social Psychology Review 9(2), 131-155. doi:10.1207/s15327957pspr0902_3 (p. 137: the IQ question “remains controversial and unresolved”).
- Christine M. Rubie-Davies and John A. Hattie (2024, online 26 August 2024; issue 2025). The powerful impact of teacher expectations: a narrative review. Journal of the Royal Society of New Zealand 55(2), 343-371. doi:10.1080/03036758.2024.2393296 (reports the Oak School result as a finding, d = 0.35).
- Janet D. Elashoff and Richard E. Snow (1971). Pygmalion Reconsidered. Charles A. Jones Publishing Company, a division of Wadsworth Publishing Company. p. 44, with Robert Rosenthal and Donald B. Rubin, Pygmalion Reaffirmed (Appendix C); and the exchange Richard E. Snow (1995), Pygmalion and Intelligence?, Current Directions in Psychological Science 4(6), 169-171, doi:10.1111/1467-8721.ep10772605, and Robert Rosenthal (1995), Critiquing Pygmalion: A 25-Year Perspective, Current Directions in Psychological Science 4(6), 171-172, doi:10.1111/1467-8721.ep10772607.
What would change it. A preregistered randomised experiment of the credible-induction kind (teachers who have not yet met the children, first and second graders, an IQ test normed for their age, testers blind to designation, raw data released) with 80% power at d = 0.2: a confident null would move the verdict toward DISSOLVED, and a replicated d near 0.2 to 0.3 toward VINDICATED for a small effect, while the large and dramatic version would stay disconfirmed either way. A new synthesis with open study-level data and an instrument-quality moderator would also move it.
VII · the check
The check
This page, recomputed against itself
The check runs when the page’s scripts have loaded the three files.
The transcription
The listing is twelve typewritten pages, each a 300 dpi scan printed sideways (PDF pages 163 to 174 of ERIC ED046892). This page read it twice. Pass A cut every page into rows and columns by the empty space between them, cut each cell into glyphs, clustered all 12,296 glyphs by shape, and gave each cluster its digit; a glyph whose nearest neighbours disagreed was left unread. Pass B was read by eye from enlarged strips of eight rows. Only ID, grade, track, designation and the twelve IQ scores were read; minority status, sex and age were never transcribed. The passes were compared cell by cell over 6,112 cells: they disagreed on 24, pass B had flagged 10 more as uncertain, and one stray mark had been read as a row. Every one was settled at the scan, at up to eight times magnification, and logged. The scan allowed two readings in 2 cells, and a printed table settled each; the page says so. The verbal May 1965 score of ID 390 reads as 60 or 69; the scan favours 69 and the book’s Table 7-3 requires it (with 60 the verbal row and mean square miss print; with 69 they match exactly). The verbal January 1965 score of ID 444 reads as 53 or 58; the book’s classroom tables for that testing, Tables A-17 and A-20, come back exactly with 58 and with no other value. A third faint cell, the reasoning May 1966 score of ID 297, read as 118, is confirmed by Table A-26. Cells still uncertain: 0. The transcription then had to pass gates it could not have been tuned to: the 382 rows, the classroom counts of Elashoff and Snow’s Table 2, every N, mean, minimum and maximum of their Tables 4 and 5 and every standard deviation there but one to its printed digit, the book’s Tables 7-1, 7-3 and 7-4, the standard deviation of all 382 pretests that Rosenthal and Rubin printed as 18.48 (the listing gives 18.48), and, for the January 1965 and May 1966 testings, the book’s appendix: 594 of the 594 numbers in its 6 tables of classroom means and standard deviations for those testings (Tables A-16 to A-18 and A-24 to A-26) come back exactly, rounded the way the book rounds, to three decimals and then to two. Its six tables of classroom gains for the same testings (Tables A-19 to A-21 and A-27 to A-29) agree as well, row by row, against a machine reading of the page images that drops too many signs to ship as a gate of its own.
Uncertainties and free choices, named
- The school-wide t: the listing gives 2.132 where the paper prints 2.15; the gap is reported, not tuned away.
- The printed 49% of unnamed children gaining 10 points against the listing’s 48.4%; the printed rho of −.94, which needs the rounded table.
- Tables 20 to 22: 103 of the 108 cells come back within 0.10 of print; the other 5, four of them adjusted for the pretest, differ by up to 0.14 for reasons the page cannot pin down (listed under the table).
- Two transcribed cells were settled by printed tables rather than by the scan alone: ID 390 by the book’s Table 7-3, ID 444 by its Tables A-17 and A-20. Both are logged with the transcription.
- One standard deviation of Elashoff and Snow’s Table 4, the reasoning pretest of the 63 first graders, comes back as 36.851 where they print 36.8, just past a rounding edge; every first-grade row was checked again at the scan, so the page takes it as their rounding. That the table prints population standard deviations for single grades and sample standard deviations for pairs of grades is the page’s reading of it.
- Every re-drawn list is drawn within classrooms. If the real assignment was also balanced by sex, the reference set is narrower than the page’s.
- Every simulated share carries Monte Carlo error, stated here at 95%: the total-IQ power 94.1% to within about 1.0 points, the reasoning-IQ power 31.0% to within about 2.0, the null rate 4.5% to within about 0.4. The seeds are fixed, so the page, the verifier and a reader’s browser get the same numbers.
- The planted sizes are the printed ones; the plant goes into the May 1965 score of one measure only and is neither rounded nor clipped.
- The page does not transcribe the TOGA raw scores (Elashoff and Snow’s Table 25); their raw-score row for grades 1 and 2, as printed, is 4.0, 6.5*, 2.5 and 4.4* and is not recomputed here.
- Free choices, each a real parameter of the engine: the treatment of extreme scores, the two norm bounds, the measure, the criterion, the grade group, one or two tails, the posttest, the grades in the claimants’ test, the contact cut, the Oak row, the Pellegrini and Hicks coding, the meta-analytic model, the planted shape and the pool, the measure and score treatment of the re-drawn lists, and the planted d and its target. The order of the listing’s rows is inert, and the verifier shows it.
Sources
- Robert Rosenthal and Lenore Jacobson (1966). Teachers’ Expectancies: Determinants of Pupils’ IQ Gains. Psychological Reports 19(1), 115 to 118. doi:10.2466/pr0.1966.19.1.115. Table 1, Table 2 and text from a retyped copy, checked against the book.
- Robert Rosenthal and Lenore Jacobson (1968). Pygmalion in the Classroom. Holt, Rinehart and Winston. Tables 7-1, 7-3 and 7-4, pp. 74 to 77.
- Janet Dixon Elashoff and Richard E. Snow (1970). A Case Study in Statistical Inference: Reconsideration of the Rosenthal-Jacobson Data on Teacher Expectancy. Technical Report No. 15, Stanford Center for Research and Development in Teaching. ERIC ED046892. Appendix B, the listing; Tables 2, 4, 5 and 20 to 22.
- Janet D. Elashoff and Richard E. Snow (1971). Pygmalion Reconsidered. Charles A. Jones Publishing Company, a division of Wadsworth Publishing Company. With Robert Rosenthal and Donald B. Rubin, Pygmalion Reaffirmed (Appendix C), and the authors’ rebuttal (Appendix D).
- Richard E. Snow (1995). Pygmalion and Intelligence? Current Directions in Psychological Science 4(6), 169 to 171. doi:10.1111/1467-8721.ep10772605. Robert Rosenthal (1995). Critiquing Pygmalion: A 25-Year Perspective. Same issue, 171 to 172. doi:10.1111/1467-8721.ep10772607.
- Stephen W. Raudenbush (1984). Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction: A synthesis of findings from 18 experiments. Journal of Educational Psychology 76(1), 85 to 97. doi:10.1037/0022-0663.76.1.85. Stephen W. Raudenbush and Anthony S. Bryk (1985). Empirical Bayes meta-analysis. Journal of Educational Statistics 10(2), 75 to 98. doi:10.3102/10769986010002075.
- Wolfgang Viechtbauer and others, metadat 1.6-0, dataset dat.raudenbush1985, GPL 2 or later.
- Lee Jussim and Kent D. Harber (2005), and Christine M. Rubie-Davies and John A. Hattie (2024), as under the verdict.
The listing is shipped as a numeric transcription beside this page, never as the scan; the terms of every source are in NOTICE.txt.