Biometrika, 1908 · “Student”, The Probable Error of a Mean, section VI

Three Thousand Pieces of Cardboard

Before he could prove the curve behind the t-test, W. S. Gosset tested it by hand. He copied the heights and finger lengths of 3000 criminals onto 3000 pieces of cardboard, shuffled them, and dealt 750 samples of four. Shuffle the same cards here. Then replay his whole experiment 100,000 times, tallied by the rule he printed, and see where his one book falls: the bad fit he blamed on coarse grouping is mostly that, and his book is still wider than 96% of honest shuffles.

The practical test

Gosset was a chemist and brewer at Guinness in Dublin, and he published as “Student”. His problem was the brewery's: how far to trust the mean of a handful of measurements when the spread has to be estimated from the same handful. He found the answer as a curve for his statistic z, the sample mean's distance from the true mean divided by the sample's own standard deviation. Before he had a proof, he checked it the only way available in 1908:

“Before I had succeeded in solving my problem analytically, I had endeavoured to do so empirically. The material used was a correlation table containing the height and left middle finger measurements of 3000 criminals, from a paper by W. R. Macdonell (Biometrika, Vol. I. p. 219). The measurements were written out on 3000 pieces of cardboard, which were then very thoroughly shuffled and drawn at random. As each card was drawn its numbers were written down in a book which thus contains the measurements of 3000 criminals in a random order. Finally each consecutive set of 4 was taken as a sample—750 in all—and the mean, standard deviation, and correlation of each sample determined.”

Student, “The Probable Error of a Mean”, Biometrika 6 (1908), p. 13.

Two things are worth knowing about the cards. The men were 3000 prisoners “undergoing their sentences in the chief prisons of England and Wales”, measured by warders on probation “as their advancement in the Service depends on the accuracy with which they measure”, and the forms “were drawn at random from the mass on the office shelves” of the Central Metric Office at New Scotland Yard (Macdonell, pp. 178 to 179). And the table Gosset used is not on the page he cites. Page 219 of Biometrika I holds Macdonell's Table VI, a separate table of 1306 criminals; the 3000 are Table III, on page 216 (R's documentation for the same data makes the same correction).

The cards, and one book of them

your book his curve his book, as printed

The heavy tails are the whole point. With only four men per sample, the sample's standard deviation is often much too small, and dividing by it throws the mean far out. Had Gosset divided by the population's true standard deviation instead, a value beyond ±1.05 would turn up in about one sample in 28 (3.6%); divided by the sample's own, his curve puts one in six (16.7%) out there. Gosset's curve for four is (2/π)(1 + z²)⁻². Today's t is his z multiplied by √(n − 1), so this is Student's t with 3 degrees of freedom in its original clothes; the change of form came later, with Fisher.

What he found, and what he blamed

Gosset tested two curves on his book: one for the 750 standard deviations, one for the 750 values of z, each on height and on finger. The z curves fitted well (Pearson's P = .56 for height, .92 for finger). The curve for the standard deviations of height did not: χ² = 48.06, P about .000,06. He did not hide it. He blamed the grouping. Scotland Yard measured heights to the nearest eighth of an inch, but Macdonell's table, and so Gosset's cards, grouped them by the inch, and the standard deviation of the population was only 2.54 inches, so four men could only produce a comb of a few hundred possible standard deviations, some of them sitting on the edges of his bins:

“As an instance of the irregularity due to grouping I may mention that there were 31 cases of standard deviations 1.30 (in terms of the grouping) which is .5117 in terms of the standard deviation of the population, and they were therefore divided over the groups .4 to .5 and .5 to .6. Had they all been counted in groups .5 to .6 χ² would have fallen to 29.85 and P would have risen to .03.”

Student (1908), p. 15.

That was an argument, not a test: in 1908 nobody could deal the book again. Here it can be dealt as often as you like, from the same 3000 cards, grouped as he grouped them.

The comb: where a sample of four can land

For heights, 16 times the square of a sample's standard deviation is always a whole number, so the standard deviations form a comb, not a spread. Four of the dozen commonest teeth stand within a hair of an edge of Gosset's bins: 1.30 inches sits at .5114 of the population's standard deviation, 1.50 at .5906, 1.785 at .7029 and 2.278 at .8967. His paper records the rule he used for such values:

“In tabling the observed frequency, values between .0125 and .0875 were included in one group, while between .0875 and .0125 they were divided over the two groups.”

Student (1908), p. 15. That is: a value within .0125 of an edge counts half to each side.

The rule matters more than it looks. Flip the buttons above and watch the same book's χ² move.

One book among a hundred thousand

We dealt his book 100,000 times: the same 3000 cards, shuffled so that every order is equally likely, dealt into 750 consecutive fours, heights by the inch and fingers in 2 mm groups, z measured from the population mean with the rare zero standard deviation sent to ±6 as he did, and everything tallied into his bins by his rule and compared with the curve values he printed. The seed is 1908; the numbers below are in replays.json, and you can deal your own thousands here.

Where his book falls

100,000 replays, seed 1908. Finger groups pair 9.4 with 9.5 cm and so on (pairing A); the other pairing (9.5 with 9.6) is in the check below. “Share as far out” is the share of replays at least as far out as Gosset's book, in the direction his value lies.
his bookprintedreplay medianmiddle 90%share as far out

What the replays show

He was right about the grouping. A population of perfectly normal men, measured exactly and tallied the same way, gives a χ² for the standard deviations of about 14 in a typical book. Gosset's inch-grouped cards, tallied by his rule, give a median of 29: most of the gap between his curve and his book is the comb, not the curve. Measured against what the grouped cards actually produce (the replays' own average tally) instead of the smooth curve, his height table gives χ² = 21.0, and 8% of replays do worse.

But his 48.06 is not a typical draw. Only 4.0% of replays reach it. The excess sits in two places: his bin .4 to .5 holds 107, and the grouping alone explains most of that spike (the replays average 87, where the curve says 64.5), but 107 is still higher than 99.3% of replays; and his bin 1.6 to 1.7 holds 11.5 where the replays average 5.6.

The same thing shows in both measurements. The standard deviation of his 750 standard deviations was .9066 for height and .9802 for finger. The replays put 3.7% of books at or above the first and 3.4% at or above the second. Because both come from the same shuffle they are slightly linked (correlation 0.19 across replays), but only 0.28% of books, about one in 360, are that wide in both. His spread of z for height (1.039) is out at 4.2% as well.

Everything else was ordinary. The z fits were typical draws (53% of replays fit height worse, 88% fit finger worse). Three samples of four men all of the same inch happened in 18% of replays. Thirty-one standard deviations of 1.30: the replays' median is 28.

A published replay, and why ours disagrees

For the centenary, James Hanley, Marilyse Julien and Erica Moodie repeated Gosset's procedure 100 times in R on the same table (The American Statistician, 2008). They found his 48.1 “just below the median (51) in our series”, which would make it an entirely ordinary book, and they concluded from the zero standard deviations that his “double precautions” of shuffling and drawing at random “appear to have worked”. Tallied into plain bins, our replays agree with them: median 54, and 68% of books do worse than Gosset's. Tallied by the half-and-half rule he printed, the median falls to 29 and his book moves into the top 4%. Their paper does not describe the half-splitting rule, and plain bins come close to their median (54 against their 51, from 100 runs), so we think that is the difference; we have not seen their code. Their z result we do not reproduce either way: they report a median z χ² of 17, where our replays give about 13 by his rule and about 15 in plain bins. Their count of samples with a zero standard deviation (21 of their 100 runs as extreme as his three) matches ours (18%).

The finger table shows why the rule cannot be skipped. In plain bins, Gosset's finger χ² of 21.80 would look suspiciously good (98% of replays worse); by his rule it is a median draw (57%). Without his rule each of his two tables looks wrong, in opposite directions. With it, only the height table stands out, and only by its spread.

What this does not show

It does not show that anything went wrong with Gosset's book. One book in 25 is as extreme as his on any one of these measures, and we looked at eight measures on each side before noticing the spread; a number chosen after looking cannot be read as a test. What survives is narrower: the one feature that stands out, a book slightly too wide, appears in both of his measurements, which a fluke in one tally would not do.

If it is not chance, we can name what would do it without being able to choose. Seven hundred and fifty standard deviations worked by hand, twice over, invite slips, and a slip usually widens a spread. A shuffle that left runs of neighbouring cards together would make some samples too alike, which also widens the spread of s while pulling its mean down; his means sit a little low (29% and 33% of replays lower), which fits, but weakly. We do not know whether his book of 3000 drawn cards survives; without it, neither reading can be tested.

None of this touches his curve. It was right, Fisher later supplied the proof Gosset lacked, and the replays here fit it as well as a grouped population can.

The check

What is ours, not his. Three choices are readings of what he wrote, not facts from it. (1) The half-splitting rule for z (“.04, .05 and .06, being divided between the two groups on either side”) we read as a hundredths digit of 4, 5 or 6; the z results barely depend on it. (2) Which millimetres he paired into 2 mm finger groups is not recorded. With the other pairing (9.5 with 9.6 cm), the finger spread is out at 1.9% instead of 3.4%, both spreads together at 0.18% instead of 0.28%, and the finger z fit at 92% instead of 88%. (3) We measured z from the table's own mean. Macdonell printed the height mean as 65.5355 inches because his classes are centred 1/16 inch above the whole inch; subtracting 65.5355 from whole-inch labels would shift every z and raise the median z χ² from about 13 to about 15.

Sources

  1. Student [W. S. Gosset], “The Probable Error of a Mean”, Biometrika 6 (1908), 1 to 25; section VI, pp. 13 to 18. Scan: archive.org/details/biometrika619081909pear; transcription: Wikisource. The tables on pp. 15, 16 and 18 were read against the page images.
  2. W. R. Macdonell, “On Criminal Anthropometry and the Identification of Criminals”, Biometrika 1 (1902), 177 to 227; the material pp. 178 to 179, Table III p. 216, Table VI p. 219. Scan: archive.org/details/biometrika119011902pear.
  3. R Core Team, crimtab, “Student's 3000 Criminals Data”, in the R datasets package (documentation), from J. R. Lobry and A.-B. Dufour's transcription.
  4. J. A. Hanley, M. Julien and E. E. M. Moodie, “Student's z, t, and s: What if Gosset had R?”, The American Statistician 62 (2008), 64 to 69. Author's copy.