The Sum Nobody Knew
In 1904 a Berlin horse named Hans seemed to do sums. His supporters reported trials in which nobody present knew the answer, and every one the society’s minutes describe came out right, 5 of 5; this page reproduces that record first. Then it hands you the test that decided the case: with nobody knowing the sum Hans got 3 of 31, and with the two men who had whispered it knowing, 29 of 31. Planted into the test’s own record at the claimed size, the ability is found in 10,000 of 10,000 doctored copies. Of the three controls the page sets against the supporters’ last reply, only one, a questioner bowing past the answer, separates it from cue reading, through a count only the 1907 German prints.
Who knows the sum?
Berlin, autumn 1904. Two men each whisper one number into a horse’s ear.
The investigators’ 31 sums (Pfungst, 1911, p. 37). A filled box is a correct answer.
The supporters’ record of trials nobody present could know the answer to, as the society’s minutes of 5 January 1905 give it (p. 470).
With the sum known to the two men who whispered it, Hans answered 29 of 31 correctly.
The record keeps totals, not trials, so the order of the boxes is a display convention: correct answers first.
Everything below is computed in your browser from two small files this page transcribed from the printed record. Numbers taken from a publication are marked and cited where they appear; every chance model, p-value and bound is the page’s own and says so. The check at the bottom recomputes the page against itself while you read it.
I · the claim, at full strength
Every trial they described came out right
On the evening of Thursday 5 January 1905 the Psychologische Gesellschaft zu Berlin met to hear its chairman, Moll, report again on “der kluge Hans”, the clever horse of Herr von Osten. In the discussion Schulrat Dr. Grabow, a retired member of the school board and one of the thirteen men who had signed the commission report of 12 September 1904, accepted that Moll had drawn the right conclusions from the material in front of him, and said the material was incomplete. He had run tests of his own, in which, in his judgement, every possible influence was excluded. The society’s minutes record the result in reported speech:
… dass Hans wirklich lesen, das Gelesene verstehen und richtig rechnen könne.
Sitzungsberichte der Psychologischen Gesellschaft zu Berlin, 5 January 1905, in Zeitschrift für pädagogische Psychologie, Pathologie und Hygiene 6 (1904), p. 470
The page’s translation that Hans could really read, understand what he had read, and calculate correctly.
The minutes do more than assert it. They print two experiments with their numbers. Grabow wrote problems on cards 11 by 4½ cm, shuffled them like playing cards and showed them to the horse “ohne dass er selbst, noch Herr von Osten, noch irgend jemand anders wusste, welcher Zettel vorn lag, weil Alle nur die Rückseite sehen konnten” (without himself, Mr. von Osten or anyone else knowing which card lay in front, because everyone could see only the back). In a second experiment Grabow and von Osten each wrote one number on a tablet that nobody else saw, and each whispered his number to the horse. Here is that record, and Karl Krall’s excerpt from a later report Grabow sent him, as printed. The last column is this page’s engine doing the arithmetic.
| record | problem | the horse tapped, as printed | the page’s arithmetic |
|---|---|---|---|
| minutes, 1905, first card | not printed as read from the card; „4 und 7“ is among the example cards | 11, then 4 for the first number and 7 for the second | 4 + 7 = 11: his three taps agree |
| minutes, 1905, three more cards | not printed | „Ebenso bei drei anderen Zetteln.“ (likewise with three other cards) | described successes; nothing to check |
| minutes, 1905, two tablets | 5 and 4, one whispered by each man | 9 hoofbeats; the tablets, turned round, read 5 and 4 | 5 + 4 = 9, as tapped |
| Krall 1912, p. 163, a card | 2 und 3 | 5, then 2, then 3 | 2 + 3 = 5, as tapped |
| Krall 1912, p. 163, a card | 12 weniger 5 | 7, then 12, then 5 | 12 − 5 = 7, as tapped |
Every trial with printed numbers adds up: 2 of 2 in the minutes, and 2 of 2 in the 1912 excerpt. In the minutes only 1 of these is a check against what was written: the tablets were turned round and read 5 and 4. For the first card the minutes print only the horse’s taps, 11 and then 4 and 7, and give „4 und 7“ among the example cards without saying that this card was turned over, so that row checks only that his three answers agree with one another. The minutes describe 5 of 5 without-knowledge trials as successes and print no failure, and they close: „Beide Versuche wurden wiederholt und gelangen.“ (Both tests were repeated and succeeded.) Taken at face value, 5 of 5 puts the horse’s success rate without knowledge at no less than 0.549, the one-sided 95% exact lower bound (the page’s computation). Counting one repetition of each experiment, the least that sentence allows, gives 7 of 7 and a bound of 0.652.
That is the claim as its authors saw it: every trial they described, with nobody present knowing the answer, came out right. The 1912 excerpt adds that the series ran over several days, „die mit wenigen Ausnahmen gelangen“ (which succeeded with few exceptions). Neither account prints how many attempts there were. Pfungst, reading the same minutes, took them to report “a large number of successful tests” (1911, p. 40); the minutes themselves describe five and say that both experiments were repeated, and the page counts no more than that.
The answers were right, and that part of the claim stands
The supporters were right that the horse’s answers were right. The investigators’ own counts, whenever the questioner knew the answer: numerals 41 of 42 (printed as 98%, the count recovered uniquely), words on placards 14 of 14, arithmetic 29 of 31 and counting on the abacus 8 of 8. Pfungst summed up these series as “90 to 100% of the responses of the various series were correct” (1911, p. 40); recomputed, the lowest of the four counted series his summary covers is 93.5%, so his summary holds. A later series, which that summary does not cover, falls just under his range: with blinders on, the horse answered 50 of 56 when he could see the questioner, 89.3% (p. 43). Hans answered, more or less readily, for about forty people (p. 31). And the commission of 12 September 1904 was right about the question it set itself: its members looked for intentional signals and found none. It also went, in Stumpf’s words, “one step beyond that which it had proposed to itself” (p. 4), adding its opinion “that unintentional signs of the kind which are at present familiar, are likewise excluded” (p. 254). Stumpf’s December report found the signals in movements so slight that they had escaped “the notice even of practised observers” (p. 262), and concluded that they need not have been given intentionally at all (p. 261).
II · the deciding control
The sum nobody knew
Carl Stumpf, the commission’s psychologist, ran a closer investigation from 13 October to 29 November 1904 with Dr. E. von Hornbostel, who kept the records, and Oskar Pfungst, a co-worker at his Psychological Institute, who conducted the experiments (Stumpf’s introduction, 1911, pp. 8 to 9). For arithmetic Pfungst used a design that removes every person who knows the answer without removing the horse’s owner:
Mr. von Osten whispered a number in the horse’s ear so that none of the persons present could hear. Thereupon I did likewise. Hans was asked to add the two. Since each of the experimenters knew only his own number, the sum, if known to anyone, could be known to Hans alone. Every such test was immediately repeated with the result known to the experimenters.
Oskar Pfungst, Clever Hans, tr. C. L. Rahn (1911), p. 37; Das Pferd des Herrn von Osten (1907), p. 32
So the claimant himself whispered half of every problem. The result, in Pfungst’s words: “In 31 tests in which the method was procedure without knowledge, 3 of the horse’s answers were correct, whereas in the 31 tests in which the method was procedure with knowledge, 29 of his responses were correct.” The German original reads „Von 62 Versuchen fielen auf die 31 unwissentlichen 3 richtige Antworten, auf die 31 wissentlichen dagegen 29.“ That is 3 of 31 against 29 of 31, a drop of 83.9 percentage points. Pfungst judged that “the three correct answers in the cases in which procedure was without knowledge evidently were accidental”. That is his judgement, not a test: he printed no chance model and no test. The panel below supplies both, and they are the page’s.
The knowledge control, as the page runs it
Pfungst printed neither the range of the addends nor a chance model, so all three rates are the page’s choices. The count needed to show ability above chance, k*, is the smallest count whose exact binomial tail is 5% or less.
At a chance rate of 10%, showing ability needs 7 of 31 correct without knowledge. Hans gave 3. Verdict: no ability seen.
The pairs Pfungst did not print
Each sum was asked twice, without and then with knowledge, but the book prints only the two totals. A paired test needs to know which unknown-sum answers were right on the same problems as the known-sum ones. It cannot be recovered, but it can be bounded: only 3 pairings fit both totals (the number of problems answered right both times is 1, 2 or 3), and the page runs the exact McNemar test on each.
| right both times | right only when known | right only when unknown | wrong both times | exact McNemar p |
|---|---|---|---|---|
| computed when the page runs | ||||
Knowledge matters under every pairing the totals allow: p lies between 2.98 × 10−8 and 8.68 × 10−7. The verdict function has three outcomes: no ability seen (fewer than k* right), ability seen, and a knower still helps, and ability seen, and no knowledge effect detected (the paired test cannot reject under at least one pairing the totals allow, which is a failure to find a difference, not proof that there is none). On the real counts it returns no ability seen at every offered chance rate (k* is 5, 7 and 11 at chance rates of 5%, 10% and 20%). The 3 correct answers are consistent with guessing for any chance rate above 0.0269, about 1 in 37, which supports “evidently accidental” under any plausible model.
The record gives no order and every control uses counts only, so reordering the rows changes nothing the page computes.
The rest of the ledger
The arithmetic series was one of eight. Here is every counted series Pfungst ran with and without knowledge before his summary on p. 40, as printed. Where the book prints a percentage, the engine recovers the count; every recovery below is unique.
| series | 1911 p. | without knowledge | with knowledge | note |
|---|---|---|---|---|
| numerals on cards nobody saw | 35 | 4 of 49 | 41 of 42 | printed as 8% and 98% |
| words on placards | 36 | 0 of 12 | 14 of 14 | with knowledge printed as 100% |
| addition, two whisperers | 37 | 3 of 31 | 29 of 31 | pairs not printed |
| counting on the abacus | 37 | 0 of 8 | 8 of 8 | the questioner’s back to the abacus |
| memory | 38 | 2 of 10 | none | one right answer was 3, the horse’s habitual number |
| calendar | 38 | 4 of 14 | none | the keeper was present and knew the four dates |
| single tones | 39 | 1 of 20 | not printed | with knowledge: “without exception” |
| compound clangs | 39 | 0 of 9 | not printed | the count 9 is in the 1907 German only (p. 33); with knowledge “all but one” |
Pfungst’s summary that “10%, at most” of the answers without knowledge were right holds for every series with a paired arm (the highest is 9.7%, the arithmetic). It does not hold for the two unpaired series, memory at 20.0% and the calendar at 28.6%, and Pfungst explains both in place: the habitual 3, and a keeper who knew the dates. The only printed table of paired trials is ten rows of one numeral series, von Osten questioning, which Pfungst gives as “an example of the course which the series tended to take”: exposed 8, tapped 14; then 8 and 8; 4 and 8; 4 and 4; 7 and 9; 7 and 7; 10 and 17; 10 and 10; 3 and 9; 3 and 3, “etc.” (p. 35). Scored, the unknown numerals come out 0 of 5 and the known ones 5 of 5.
III · the control on the control
Could the test have found a horse that could add?
A test that could not have found the ability proves nothing by failing to find it. So the page gives the claim its size and plants it. The 31 unknown-sum answers become 31 rows, 3 right and the rest wrong. The claim is that Hans computes the sum himself, so in a copy of those rows each recorded failure independently becomes a success with probability a, the chance that on that trial he added the numbers. The claimed size is the ability at which the doctored unknown-sum arm is expected to match the known-sum arm of the same session, the claim that knowledge makes no difference: a = 26/28 = 0.929, read from the frozen counts. Each doctored copy then goes through the same unmodified function that produced the real verdict, against the real known-sum arm.
At the claimed size, in 10,000 doctored copies (seed 11), the control read “ability seen” in 10,000 of 10,000 and “ability seen, and no knowledge effect detected” in 10,000 of 10,000. Exactly, running the control on every possible doctored count weighted by its probability, “ability seen” has probability 1.000000 (at least 1.000000 at every offered chance rate) and “no knowledge effect detected” 0.999986. The control could have confirmed the claim. The page’s rule for that sentence: “ability seen” with probability at least 0.95 at every offered chance rate, so that the control would miss a real ability no more often than its tests raise a false alarm; below that line the sentence would read that the control could not have confirmed the claim, and its null would be inconclusive. It would have seen even the weakest rate the supporters’ own record allows: planted at a = 0.501, the ability that gives an expected rate of 0.549, it reads “ability seen” with probability 0.99999, and “a knower still helps” with probability 0.955, because that rate is still below the known-sum arm. With 80% power the 31 trials would see an ability as small as 0.103, 0.188 and 0.344 at chance rates of 5%, 10% and 20%.
Plant an ability and rerun the unmodified control
Marked values: the claimed size, the supporters’ weakest face-value rate, and zero, which must return the real verdict.
Planted at the claimed size, the control read “ability seen” in 10,000 of 10,000 doctored copies.
One doctored copy. Empty: the failures the record kept. Hatched: failures the plant turned into successes. Filled: the 3 recorded successes.
Solid line: probability the control reads “ability seen”. Dashed: “ability seen, and no knowledge effect detected”. Both are exact, computed by running the control on every doctored count.
Grade A, with a stated limit: the plant can only go into what the record preserves, 31 outcomes and their count. It cannot plant into the addends, the rhythm of the tapping or the pairing of the repeats, none of which were printed.
The second control: a knower the horse cannot see
Pfungst then asked by which sense the horse got his cue. Von Osten believed the horse heard the question spoken inwardly. Pfungst put blinders on the horse; the questioner, who always knew the answer, stood to the right, where the horse knew him to be present and could hear him but not see him (p. 42). Hans would “always make the most strenuous efforts to get a view of the questioner” (p. 43), so in some tests nobody could tell whether he had managed it, and Pfungst kept a third category, undecided. With blinders 15 by 15 cm: not seen 2 of 35, seen 50 of 56, undecided 2 of 11 (p. 43; the counts recovered uniquely from the printed 6%, which Pfungst says was only two, 89% and 18%). The page’s one-sided Fisher test that unseen is worse than seen gives p = 2.37 × 10−16.
Planted the same way, an ability to answer unseen at the seen rate (a = 0.886, from the frozen counts) makes the unmodified sight test read “no sight effect detected” with exact probability 0.997 (9,979 of 10,000 doctored copies, seed 3). This control, too, could have confirmed a horse that did not need to see anyone.
Von Osten’s own blinders were smaller, and with them a third of the tests were undecided: not seen 6 of 25, seen 36 of 44, undecided 28 of 39 (pp. 44 to 45). Counting the undecided with the unseen, as von Osten had done, Pfungst wrote, “then one would have been led to the conclusion that the horse did not need visual signs” (p. 45). The page lets you make that choice. Counted his way the unseen horse answers 34 of 64, which is 53.1% and looks like a horse who can answer blind; the Fisher test still finds the unseen answers worse than the seen ones, p = 0.0018. An observer who counted rates and did not compare arms would have kept the claim.
The sight control
Large blinders, undecided set aside: unseen 2 of 35, seen 50 of 56; p = 2.37 × 10−16, sight matters.
The four cells, and the one nobody ran
| questioner seen | questioner unseen | |
|---|---|---|
| answer known | known and seen: every with-knowledge arm | the blinder tests |
| answer unknown | every without-knowledge arm |
IV · the reply the control could not answer
They predicted the failure
The supporters had an explanation for failure before the failures were counted. At the commission’s session of 12 September 1904, when a test was proposed in which von Osten would not know the number, he “said that he thought that this method was somewhat risky, since the horse would be aware that he, Mr. von Osten, did not know the number, and might therefore be in a humor to play some prank” (the commission’s records as abstracted by Stumpf, 1911, p. 258). Pfungst’s first chapter reports the same belief among the horse’s friends (p. 24). Grabow restated it in the minutes of 5 January 1905: the commission had not taken into account the psyche of ‘clever Hans’,
der nur dann richtige Antworten gebe, wenn er merke, dass die Richtigkeit der Antworten kontroliert [sic] werden könne, natürlich nach seinem Erachten
Zeitschrift für pädagogische Psychologie 6, p. 470 (the print spells kontroliert with one l)
The page’s translation who gives correct answers only when he notices that their correctness can be checked, in his own estimation, of course.
Karl Krall, who worked with Hans and trained horses of his own, built a chapter on it in 1912. Failure without knowledge proves nothing, he argued, because „das Pferd versagt sofort, wo es sich dies gestatten zu können glaubt“ (the horse fails at once wherever it believes it can allow itself to; Denkende Tiere, p. 162), and on his own experience „Hans war sehr wohl imstande, ‚unwissentlich‘ dargebotene Aufgaben zu lösen“ (Hans was perfectly able to solve problems presented without knowledge; p. 164). One former supporter went the other way. Schillings, who had exhibited the horse through the summer, told the same meeting in 1905 that the horse might react to other bystanders when the questioner did not know the answer, and ended: „Es bleibt aber keine andere Erklärung.“ (But no other explanation remains; p. 471.)
A reply that predicts the result of a control cannot be refuted by that control. The 3 of 31 is exactly what a withholding horse would produce. So the page takes the reply at full strength, in both its printed forms, and asks which of the investigators’ controls could tell it apart from cue reading. What each explanation predicts is the page’s reading of each source; press a control to take it out of the reckoning.
Four explanations against three controls
| explanation | nobody knows | knower unseen | knower stays bent |
|---|---|---|---|
| H · Hans calculates | contradictedpredicts: answers correctly | contradictedpredicts: answers correctly | contradictedpredicts: taps the sum |
| W1 · withholds when nobody can check | not contradictedpredicts: may fail | contradictedpredicts: answers correctly | contradictedpredicts: taps the sum |
| W2 · withholds wherever he believes he may | not contradictedpredicts: may fail | not contradictedpredicts: may fail | contradictedpredicts: taps the sum |
| C · reads a knower’s involuntary movements | not contradictedpredicts: fails | not contradictedpredicts: fails | not contradictedpredicts: taps where the bow ends |
| observed | 3 of 31 (exact 95% interval 2.0% to 25.8%) right | 2 of 35 (exact 95% interval 0.7% to 19.2%) right | 26 of 26 (exact 95% interval 86.8% to 100.0%) ended where the bow ended |
Not contradicted by any included control: C (He reads the involuntary movements of someone who knows).
The rule: a cell is marked contradicted only when the whole exact 95% interval of the observed rate lies on the wrong side of the line its prediction draws, 50% for “answers correctly” (or “taps where the bow ends”) and 25% for “fails” (or “taps the sum”). “May fail” predicts nothing and is never contradicted. The blinder and edition settings are shared with their own panels.
Why each explanation predicts what it does (the page’s reading)
With all three controls in, the knowledge control contradicts only H, the claim that Hans calculates; the sight control contradicts H and W1; and the posture test contradicts H, W1 and W2. So W2, Krall’s reply, falls to P alone, the posture test, and the one explanation left standing is C. Count the small blinders von Osten’s way and the sight column turns against the cue explanation instead, leaving none of the four standing, which is the sign that the counting, not the horse, has changed. Add the 1911 wording to that counting and the survivors are W1 and W2, the supporters’ two replies: the one setting in which the matrix leaves exactly their explanations standing is the one that scores 39 tests in which the horse may have seen the questioner as tests in which he could not.
The test that separates withholding from cue reading
Withholding can explain a wrong number. It cannot say which wrong number. Pfungst found that the horse started tapping when the questioner leaned forward, even when only the head was lowered (29 of 30 such tests, p. 57), and stopped when the questioner straightened. So he asked a question and then deliberately stayed bent past the answer:
If, for instance, I stooped forward after having told the horse to tap 13, and if I purposely remained in this position until I had counted 20, he would, without any hesitation, tap 20. If I asked him to add 3 and 4, but did not move until 14 was reached, he would tap 14.
Pfungst 1911, p. 57; 1907, p. 46
The two printed examples, replayed
Press a button to replay a printed example. The replay shows what the book reports; it is not a model of the horse.
1907, p. 46: „26 derartige Versuche hatten alle den nämlichen Erfolg.“ 26 of 26, the lower end of its exact two-sided 95% interval being 0.868.
Here the translation matters. The 1911 English says “Twenty-six such tests gave similar results”, which prints no count, and a page reading only the English would have to call the decisive series uncounted. The 1907 German says „26 derartige Versuche hatten alle den nämlichen Erfolg“: twenty-six such tests all had the same result, which the page reads as 26 of 26 ending where the questioner straightened. Switch the edition to English and the matrix can no longer mark the posture column, and W2 and C both survive. The whole separation of withholding from cue reading rests on the word „alle“.
Three further printed series point the same way, and the page shows them as corroboration rather than putting them in the matrix, so the matrix is the page’s selection of three controls, not every test that bears on the reply. With Stumpf and Pfungst each thinking of a number beside the horse, “In ten such tests Hans always tapped my number” (10 of 10, p. 72). And when the questioner bent further forward halfway through, the tapping sped up in 31 of 34 tests (p. 64). And when a group of twenty onlookers knew the answers and only the questioner did not, the horse answered 2 of 21 problems correctly (p. 69). Stumpf’s report states the posture result in general terms, without a count: Pfungst “can now call forth at will all the various reactions of the horse by making the proper kind of voluntary movements, without asking the relevant question” (p. 262).
V · two records of one design
Grabow’s tablets and Pfungst’s whispers
Grabow’s tablet test, two people each whispering one number that the other did not know, is the design Pfungst used for his 31 sums. Taken at face value and treated as independent trials with one common success rate, the supporters’ 5 of 5 and the investigators’ 3 of 31 are hard to reconcile with one shared rate: the page’s one-sided Fisher test gives p = 1.49 × 10−4, or 9.51 × 10−6 counting one repetition of each experiment. So the two records disagree, and something about the conditions or the counting differed. The page does not say what.
Two records, one design
Supporters 5 of 5 (face-value lower bound 0.549) against the investigators’ 3 of 31: one-sided Fisher p = 1.49 × 10−4.
„Beide Versuche wurden wiederholt und gelangen“ says the experiments were repeated, not how often. One repetition each is the least it allows; the page counts no more than that, and no repetition of a record that fails its own anchor.
What it can say is what each record prints. The minutes give no number of attempts, no failures, where each man stood, or whether each whisper could be heard by the other. Pfungst made the same objection in 1907: “A thorough analysis of his experiments was not possible, because the conditions under which they were conducted were not adequately specified” (1911, p. 40), and added his own judgement, that he had “no doubt that the successful responses of the horse were due solely to the absence of precautionary measures”. That is his judgement. Stumpf’s report had already set the terms: “statements based solely upon memory, without specific report of experimental conditions, prove nothing” (p. 265). Pfungst’s record prints its denominators and its conditions; the supporters’ record prints no denominator and only part of its conditions, and the 1912 excerpt concedes “few exceptions” without counting them. Only one of the two gives a reader the numbers to count with.
VI · the supporters’ method, run on nothing
A responder who could not know
The supporters’ evidence had one shape: someone who knows the answer asks, the responder answers, and correct answers are counted as the responder’s knowledge. Two historical runs of that procedure were made where the counted effect cannot have been what it would be counted as: once with a responder who could not have known the answer by any faculty of his own, and once for a factor that, in other hands, made no difference.
From November 1904 Pfungst stood in for the horse in the Berlin laboratory (p. 102). Subjects fixed their minds on one of four or six concepts; Pfungst, facing them, named the one he thought they had in mind. “With twelve subjects (a total of 350 tests) I made an average of 73% correct responses” (p. 106). If that average is taken over all 350 tests, the correct count is one of 4 values, 254 to 257, and the page prints the range rather than choosing. Guessing one of four gives an expected 87.5 (one of six, 58.3), and the chance of 254 or more by guessing one of four is 1.31 × 10−77 (the page’s computation; the blanks subjects were told to mix in would only lower these chance rates). Counted the supporters’ way, and given the same face-value bound the page gives their 5 of 5, Pfungst’s rate at reading minds is at least 0.684 to 0.693 (depending on the count), higher than the 0.549 the supporters’ record earns. He read movements: “Very slight involuntary movements of the head and eyes” (p. 106). The German says „durchschnittlich 73%“, which could also be the mean of the twelve subjects’ rates; the recovery above assumes the pooled rate.
Von Osten ran the procedure himself on a factor of his own choosing. He believed Hans heard the question spoken inwardly, and showed it: with the sound supposedly blocked, 20 tests of which “95% of the responses were incorrect”, that is 1 of 20 right; unblocked, “not a single answer was wrong” in 28, 28 of 28 (p. 42). That is a contrast of 95.0 percentage points, one-sided Fisher p = 1.73 × 10−12, larger than the knowledge effect of 83.9 points. Pfungst repeated both kinds of test and “always found some correct responses” when the horse supposedly could not hear, and got correct answers without whispering the question at all (p. 42). He printed no counts for that repetition, so the claim that the factor did nothing rests on it and on the silent posture tests. A questioner who knows which condition is running produced a contrast as large as the one this page is about, for a factor that, in other hands, made no difference.
VII · the verdict, dated
What the record says, as of now
With nobody knowing the sum Hans answered 3 of 31, against 29 of 31 when it was known; fed copies of its own record with the claimed ability planted in, the same test found it in 10,000 of 10,000; and of the three controls this page tests, only the posture test, 26 of 26 in the 1907 German, contradicts the supporters’ last reply, Krall’s, that the horse withheld answers wherever he believed he could.
ARTEFACT
As of . Scope: the claim that Hans himself read and calculated, so that his correct answers did not depend on anyone present knowing the answer. The correct answers were real, and so was the horse’s perception of human movement; the arithmetic was not his. Decided by contamination: the answer reached the horse from the people who knew it, and the deciding tests removed that knowledge, removed the sight line, and then planted the signal on purpose. From the commission report of 12 September 1904 to Stumpf’s report of 9 December 1904 took 88 days; Pfungst’s full account followed in 1907.
Not withdrawn: Grabow restated the claim in January 1905 and Krall defended it in 1912, and this page located no withdrawal by von Osten. One early supporter did abandon it. Schillings, who had exhibited the horse, was led by the first tests without knowledge “to replace his hypothesis of independent conceptual thinking by one of some kind of suggestion” (Stumpf’s introduction, p. 9), and when told the investigators’ conclusion “embraced it without wavering” (p. 265).
What would change it. Hans cannot be retested, so only the record can move this verdict: a counted without-knowledge protocol from 1904 or later, with every attempt and failure logged day by day and no knower in the horse’s sight, showing correct answers well above any chance model; or the original records kept by von Hornbostel, if they were located and contradicted the printed totals. This page has located neither.
Sources for the verdict
- Carl Stumpf, report of 9 December 1904, printed in Oskar Pfungst, Das Pferd des Herrn von Osten (Leipzig: Barth, 1907), Beilage IV, from p. 185, and in the English translation, Clever Hans (New York: Holt, 1911), Supplement IV, pp. 261 to 265. “The horse failed in his responses whenever the solution of the problem that was given him was unknown to any of those present.” “The horse failed again whenever he was prevented by means of sufficiently large blinders from seeing the persons, and especially the questioner, to whom the solution was known.” (p. 261)
- Oskar Pfungst, Das Pferd des Herrn von Osten (Der kluge Hans), Leipzig: Johann Ambrosius Barth, 1907, chapter 2; English translation by Carl L. Rahn, Clever Hans (The Horse of Mr. von Osten), New York: Henry Holt and Company, 1911, chapter II. Project Gutenberg text; scans at the Internet Archive: 1907, 1911.
- Laasya Samhita and Hans J. Gross, The “Clever Hans Phenomenon” revisited, Communicative & Integrative Biology 6(6): e27122, November 2013, doi:10.4161/cib.27122. Cited for the verdict. Its abstract places the signals in the questioner’s face; Pfungst placed the stopping signal in a movement of the head, “not ... from the face of the questioner” (1911, p. 62), while noting that raising the eyebrows or dilating the nostrils “seemed also to be efficacious”, perhaps with a slight upward movement of the head he could not rule out (p. 63). This page follows Pfungst.
- Matyáš Moravec, How a ‘psychic’ horse named Clever Hans changed the way we do science, The Conversation, republished by Phys.org on 27 August 2026: “Although it turned out that these horses and dogs had no such mathematical powers, the studies conducted by psychical researchers helped uncover new knowledge.” The most recent dated public account this page located.
Stumpf’s report was careful about the man as well as the horse: “No one has the right, however, to charge an old man, who has never had a blemish on his reputation, with having invented a most refined network of lies, if the facts can be explained in a satisfactory manner in some other rational way” (pp. 263 to 264). The signals were given without the questioners knowing they gave them; that is the whole finding. And he called the horse’s perception of those minimal movements “astounding in the highest degree” (p. 262).
One historian has put forward the hypothesis that Stumpf wrote or edited substantial parts of Pfungst’s book, and argues that the book served to conceal that Stumpf had for a long time assumed higher mental gifts in the horse (Horst Gundlach, Psychologische Rundschau 57(2), 96 to 105, 2006; the page read the abstract only). The argument concerns authorship and motive, not the counts, but it is one reason this page names whose record each count comes from.
The timeline
- Von Osten begins teaching the horse; Stumpf speaks in December 1904 of “four long years” (p. 264).
- Schillings first comes to the courtyard, and soon exhibits the horse himself (p. 2).
- Thirteen signatories, Stumpf and Grabow among them, find no intentional signs under the conditions they kept, and judge the unintentional signs then familiar excluded too (pp. 253 to 254).
- Schillings and Pfungst make the first tests in which nobody present knows the answer (pp. 8 to 9).
- Stumpf, von Hornbostel and Pfungst’s investigation in the courtyard (p. 9).
- Laboratory tests begin, Pfungst in the horse’s place (p. 102).
- Stumpf’s report.
- Grabow’s remarks to the Psychologische Gesellschaft (p. 470).
- Pfungst’s book, in German; , in English.
- Krall’s Denkende Tiere replies.
VIII · the check
The check
Recomputed in your browser, now
The live check runs when the page’s scripts load.
What this page rests on, and what it chose
- Transcriptions. Both frozen files were typed by this page from the page images: the 1905 minutes (leaves 479 and 480 of the Internet Archive scan), Krall 1912 p. 163 (leaf 186), and the 1907 German pp. 32, 33, 45 and 46; the 1911 English was read through the Project Gutenberg page anchors and checked against the running heads of the 1911 scan. Every entry carries its printed words.
- Recovered counts. Numerals 8% and 98%, words 100%, blinders 89%, 18%, 24%, 82% and 72%, and von Osten’s 95% each round from exactly one count. The laboratory’s 73% of 350 does not; the page gives the range.
- Choices, each yours to change. The chance model (three rates); whether the minuted repetitions count; whether undecided blinder tests are set aside or counted as unseen; which blinders; which edition’s words for the held bend; the planted ability and its seed; the row order, which is inert and shown to be.
- The page’s own readings. What each explanation predicts for each control; the 50% and 25% lines; the one-sided tests at 5%. Pfungst printed no test of any kind.
- What it refuses. Which sums the 31 trials were, their order, or the pairing of each repeat; a single paired p-value; a rate for any series whose count was never printed (tones and clangs with knowledge, Schillings’s small-blinder series, the held bend in the English, Grabow’s repetitions); a data cell for a knower who is both absent and unseen; a planted ability outside 0 to 1 or a plant into the known-sum arm.
- Not located. The original records kept by von Hornbostel; the individual sums; any day-by-day protocol of Grabow’s tests. The page says not located, not lost.
The further result, and how far it goes
The citable result is the matrix: of the three controls the page sets against it, only the held-bend posture test contradicts the supporters’ reply that the horse withheld answers wherever he believed he could, and only through the count the 1907 German prints; the other posture series above point the same way but are not in the matrix. On what exists already: we searched the web through a general search engine, Europe PMC, the Internet Archive catalogue and the Wasteland’s own index on 2026-09-22 and did not find an interactive page or published analysis that sets Grabow’s 1905 without-knowledge record beside Pfungst’s 1904 counts, or that tests each printed reply of the claimants against each of Pfungst’s controls.