At full strength / psychology / a claim, its replication, and what the replication could see

The Hours Between Best and Good

In 1993 a study of Berlin violinists reported that the best students had practised alone for 7,410 hours on average by 18 and the good ones for 5,301, and called the match between skill and practice complete; the 10,000-hour rule grew out of it. A 2019 replication released its 39 practice histories. Run the 1993 test on them, then plant the 1993 gap into the 2019 best violinists: at its printed size the replication had about a one-in-four chance of confirming it, so whether a conservatory's best student violinists had accumulated more solitary practice by 18 than its good ones is still open.

Two conservatories, one question: did the best students practise more than the good ones?

Accumulated practice alone by age 18. Left, the three group means the 1993 paper printed (7,410, 5,301 and 3,420 hours; no individual values were ever published). Right, the replication's 39 violinists, one dot each, once you run them. After the plant, the orange dots are the planted copy of the best group and the gold line marks the real best mean.

The 2019 side is empty until you run it.

Everything below is computed in your browser from two small files: the numbers the 1993 paper printed, and the 39 practice histories the 2019 replication released. A number taken from a paper is marked as printed and says where. Nothing you do here is sent anywhere.

Both studies measured the same thing in the same way: adults estimating, year by year, how many hours a week they had practised alone as children. Neither study watched anyone practise. Every number on this page inherits that limit, and inherits it equally.

The page has not recomputed its figures in this browser yet.

I / The claim, at the strength it was printed

Berlin: ten soloists, ten good violinists, ten future teachers

In 1993 K. Anders Ericsson, Ralf Th. Krampe and Clemens Tesch-Römer published in Psychological Review a study of violin students at the Music Academy of West Berlin. The professors nominated students with the potential for careers as international soloists; of 14 nominated, 10 took part, and the paper calls them the best violinists. From a larger pool of good students in the same department the authors chose 10, matched on sex and age, and a further 10 came from the academy's music-education department, which had lower admission standards. The paper calls that group the music teachers, because teaching was their most likely future. Each violinist estimated, for every year since they began, how many hours a week they had practised alone. The weekly estimates of practice alone can easily be converted to estimated yearly amounts by multiplication of the number of weeks in a year.

The paper's test was made at 18, and it says why: To avoid any confounding influences from the activities at the music academy, we statistically analyzed the amount of practice the young violinists had accumulated by age 18. Here is the result, in the authors' words.

At this age, the best young violinists had accumulated an average of 7,410 hr of practice, which is reliably different from 5,301 hr, the average number of hours accumulated by the good violinists, F(1, 27) = 4.59, p < .05. The average of the best two groups was reliably different from that of the music teachers, who had accumulated 3,420 hr of practice by age 18, F(1, 27) = 11.86, p < .01. Hence, there is complete correspondence between the skill level of the groups and their average accumulation of practice time alone with the violin.Ericsson, Krampe and Tesch-Römer (1993), Psychological Review 100(3), p. 379.

And the claim was larger than one comparison. The abstract of the same paper:

Individual differences, even among elite performers, are closely related to assessed amounts of deliberate practice. Many characteristics once believed to reflect innate talent are actually the result of intense practice extended for a minimum of 10 years.Ericsson, Krampe and Tesch-Römer (1993), abstract, p. 363.

The claim, recomputed here from the only numbers ever printed

best, good, teachers7,410 / 5,301 / 3,420mean hours by 18, printed
pooled SD from F = 4.592,201.2 hEricsson (2014): 2201
second F, predicted11.857printed 11.86; rounding allows 11.828 to 11.885
best over goodd = 0.958p = 0.041 for F(1, 27) = 4.59

The paper printed two F tests and no standard deviations. Both tests divide by the same pooled variance, so the first test fixes that variance (4,845,186.27, a standard deviation of 2,201.2 hours) and the variance then predicts the second test. Predicted 11.857; printed 11.86, which lies inside the interval (11.828 to 11.885) that the printed rounding allows. The two published tests are one consistent set of numbers, and the claim stands at full strength: the best violinists ahead of the good by 2,109 hours, about 0.96 pooled standard deviations, p = 0.041; the soloist-track students ahead of the future teachers at p = 0.0019.

Try a wrong number, and watch the two tests stop agreeing

With the printed numbers the predicted second F is 11.857. With the teachers' mean at 3,520 it would be 11.06, outside the interval: the check can fail.

What the check can and cannot see: it asks whether the two printed F values fit one pooled variance. Add the same number of hours to all three means, or multiply all three by the same factor, and both F values stay as they were, so the check still reads inside. It vouches for the two tests agreeing with each other, not for the hours themselves.

II / Where the 10,000 comes from

Not from the sentence above

The test was made at 18, and the number was 7,410. The paper's Figure 9 draws accumulated practice on to age 20, and there the best group's curve ends close to 10,000 hours (read by eye from the printed figure; this page has not digitised it). In the prose, 10,000 hours appears once, in a passage about how the body adapts to training: the effects of over 10,000 h of deliberate practice extended over more than a decade (pp. 393 to 394).

Malcolm Gladwell's 2008 book Outliers made the number a rule. Ericsson and Harwell (2019) record that he called it the magic number for true expertise: ten thousand hours. Their answer: Although our research showed that an extended period of training and practice was required for attaining international-level performance, there was no evidence for a magical number. The number kept travelling anyway. When Ericsson died in June 2020, his university's notice said that he pioneered the concept that it takes 10,000 hours of practice to become an expert.

The replication reproduces the round number on its own terms: By age 20, both the best and good violinists had accumulated more than 10 000 h of practice alone on average. Recomputed from the released histories, the best group averaged 10,586 hours by 20, all 13 histories complete. The good group's 11,804 counts 7 histories that stop before 20, so it is a floor; the 6 good violinists who had reached 20 averaged 9,191.

This page does not grade the 10,000-hour rule. It grades a narrower sentence, and the sentence is the one the 1993 authors wrote.

III / The replication, with its rows in the open

Cleveland: the same design, run blind

Brooke N. Macnamara and Megha Maitra ran the study again at the Cleveland Institute of Music. Faculty nominated 24 students with the potential for careers as international soloists, and 13 agreed to take part: the best violinists. 13 good violinists were matched to them on sex and age as far as possible, and 13 violin students from the music department of neighbouring Case Western Reserve University formed the least accomplished group. Recruitment ran from the summer of 2014 to the autumn of 2018. The analysis was preregistered in April 2017, before the first author had seen the data: The first author did not collect data from any of the student violinists and has not coded or looked at the data.

The interviews were double-blind: interviewers did not know a violinist's group, and violinists were not told there were groups. The replication's authors observe: There is no indication in Ericsson et al. that experimenters were blind to the participants' skill level. This page reports that as a difference of design and nothing more. Here is what the replication found, in its authors' words.

We found a significant effect of group for accumulated practice alone until age 18, χ²₂ = 13.90, p = 0.001, η² = 0.26. However, the best violinists (M = 8224, 95% CI [6400, 10 048], range = 3978–14 664) had not accumulated significantly more practice alone by age 18 than the good violinists (M = 9844, 95% CI [6937, 12 751], range = 3120–21 268), t₂₄ = −0.93, p = 0.364, d = −0.38. The good violinists had accumulated significantly more practice alone by age 18 than the less accomplished violinists (M = 4558, 95% CI [3264, 5851], range = 2522–10 972), t₁₆.₅₇ = 3.26, p = 0.005, d = 1.33.Macnamara and Maitra (2019), Royal Society Open Science 6: 190327, p. 15. CC BY 4.0.
We did not replicate Ericsson et al.'s [1] major result of ‘complete correspondence between the skill level of the groups and their average accumulation of practice time alone with the violin’ (p. 379).Macnamara and Maitra (2019), p. 16.

The 2019 rows, recomputed from every weekly estimate

Practice measure
Test for best against good
One dot per violinist. The line is the group mean; the band is the mean plus or minus 1.96 standard errors, the interval the paper printed.

best mean 8,224 h (SD 3,355; interval 6,400 to 10,048; range 3,978 to 14,664)

good mean 9,844 h (SD 5,348; interval 6,937 to 12,751; range 3,120 to 21,268)

least accomplished mean 4,558 h (SD 2,379; interval 3,264 to 5,851; range 2,522 to 10,972)

best minus good = −1,620 h. Levene F = 2.68, p = 0.115.

Student t(24) = −0.93, p = 0.364; d = −0.38 by the paper's 2t/√df, −0.36 with the pooled SD

Kruskal-Wallis H = 13.90, p = 0.00096; eta squared 0.259

Accumulated practice by age for the three 2019 groups, counting only the histories that reach each age; the open diamonds at 18 are the 1993 printed means, which are practice alone, so they are drawn only against the practice-alone curves. Past 18 the curves go on but no test runs there, and they are drawn from fewer violinists: complete histories at 19, best 13, good 12 and least accomplished 10; at 20, 13, 6 and 8.

At 18, with practice alone, every number the replication printed comes back from the weekly rows (the full list is in the check below): best 8,224 h, good 9,844 h, least accomplished 4,558 h; best against good Student t(24) = −0.93, p = 0.364; good against least accomplished Welch t(16.57) = 3.26, p = 0.005. And the replication's plainest sentence is true of its own rows: In fact, the majority of the best violinists had accumulated less practice alone than the average amount of the good violinists. Counted here: 8 of 13.

Two conventions matter for that match. The printed intervals are the mean plus or minus 1.96 standard errors (a t interval does not reproduce them), and the printed d is 2t divided by the square root of the degrees of freedom; with the pooled standard deviation instead, best against good is d = −0.36. The 2019 preregistration names the rule for choosing between the two t tests: if Levene's test were to be violated for a t-test, we would conduct a Welch's t-test. Levene's test, read here with deviations from the group mean, picks Student for best against good (p = 0.115) and Welch for good against least accomplished (p = 0.0073), which are the two tests the paper printed.

Move the age below 18 and every mean and test moves with it. Move it past 18 and the page refuses to test: at 19, 1 good and 3 least accomplished violinists had been interviewed at 18, and the workbook holds no estimate for years they had not yet lived. Ask for the professionals and the page refuses too. The replication interviewed 4 and ran no test on them; its supplement says: Due to the small size, we do not conduct any inferential statistics with this group. This page does not ship their rows. The 1993 study interviewed 10 and tested them once, against its best students at 18, as a further test of the relation between performance and practice: The average for middle-aged violinists is 7,336 hr, which is so close to the average of 7,410 hr for the best young violinists, that the difference is not statistically significant. (p. 380). That test needs the 1993 rows, which were never released, so this page cannot rerun it.

IV / The 1993 test, on the 2019 rows

One step comes back at nearly full strength; the other comes back reversed

The 1993 paper describes its analysis in one sentence: the hypothesized differences between the three groups of violinists are represented by two orthogonal contrasts (p. 374). That is one error term shared by two comparisons: best against good, and the average of the best and good (the paper calls them the soloist students) against the teachers. The replication describes the same analysis as two ANOVAs; the page uses the paper's own description.

Run exactly that procedure on the 39 released histories. Best and good against least accomplished comes back F(1, 36) = 11.45, p = 0.0017, against the printed 11.86. Best against good comes back F = 1.12 (p = 0.30) with the sign reversed: the good group ahead, by 1,620 hours. The first half of the claim, that students on a soloist track had practised far more than students heading for teaching, came back at nearly the printed strength, with the soloist-track students ahead by 4,476 hours. The second half, the best over the good, did not.

V / The control on the control

Could this replication have said yes?

A replication that finds nothing can only speak against what it could have found. So before reading the replication's null result as an answer, the page asks what it would have reported if the 1993 gap had been there. The replication did publish a power statement:

Our sample size is large enough to detect an effect size of η² = 0.48, which is the effect size found by Ericsson et al. [1], with greater than 99.9% power.Macnamara and Maitra (2019), p. 5.

That is the sensitivity of the three-group test, and the replication passed that test (Kruskal-Wallis chi-squared 13.90, p = 0.001). The step it reports as not replicated is best against good, and for that step it printed no sensitivity at all. So the page measures one, directly: it plants the 1993 gap into a copy of the replication's own weekly histories and runs the replication's own, unmodified analysis on the copy.

Grade A, by injection. The claimed effect goes into the control's own released data, and the control's own analysis runs again on the copy. The check trusts only that the workbook records what the violinists said. It does not lean on any sensitivity figure the replication printed, because the one it printed belongs to a different test.

Plant the claim into the control, then run the control

What "the claimed size" means
What to plant

Planted: 2,109 h, hours as printed. Each of the 13 best violinists gains 3,729 h, spread over their weekly estimates from their first practising year to 18; good and least accomplished rows untouched.

best against good: Student t(24) = 1.20, p = 0.240, not confirmed

good against least accomplished: Welch t(16.57) = 3.26, p = 0.0048, confirmed

three groups: Kruskal-Wallis p = 0.00008, eta squared 0.409

2,489 of 10,000 resampled replications confirm best over good: power 0.249 (Monte Carlo standard error 0.004; seed 20190821)

The replication's own estimate, −1,620 h, tested against a true gap of 2,109 h: t(24) = −2.13, p = 0.044

At this reading the replication was unlikely to confirm the claim: the copy carrying exactly this gap comes back at p = 0.240, and 25% of resampled replications confirm it, so its non-significance is inconclusive at this size.

Each bar counts resampled replications: 13 best and 13 good violinists drawn with replacement from the planted copy, run through the replication's analysis. Green: the analysis confirmed best over good. Grey: it did not. The dashed line is the planted gap; the blue line is what the replication actually measured.

At the default reading the answer is: probably not. Planted at the printed 2,109 hours, the gap comes back from the replication's own test at Student t = 1.20, p = 0.240, and only 2,489 of 10,000 resampled replications confirm it (power 0.249). The reason is spread: the 2019 best and good groups have standard deviations of 3,355 and 5,348 hours, against the 2,201 implied by the 1993 tests, so the printed gap is only d = 0.472 in the 2019 spread. A normal-theory check agrees: 0.21 at d = 0.472 with 13 a group. Power is a chance, not a yes or no, so the page's words follow a stated rule: The words follow bootstrap power: below one half the replication was unlikely to confirm the planted gap, from one half to 80% it could have, and from 80% it probably would have.

Read the claimed size without its units and the answer changes, barely. As the printed ratio of means the gap is 3,916 hours in the 2019 good group's terms, and the planted test gives p = 0.035 (power 0.595); as the printed standardized gap, d = 0.958 in the 2019 three-group spread of 3,895 hours, it is 3,732 hours, p = 0.043 (power 0.559; the normal-theory figure for this planted gap, d = 0.836 in the best-and-good spread the replication's test uses, is 0.53). By that rule the replication could have confirmed the claim at either reading, and at both it falls short of the 80% chance a replication is designed for. All three readings are live above, and the default is the one least favourable to the replication. Planting the whole 1993 pattern changes nothing about the first step (p = 0.240) and costs the second: good against least accomplished falls to p = 0.263, while the three-group test still passes (p = 0.025). That plant also drives 1 total and 65 weekly estimates below zero, which is why it is arithmetic and not a possible history.

And the test that was unlikely to say yes did say something. Test the replication's own estimate, −1,620 hours, against a true gap of 2,109: t = −2.13, p = 0.044. Once the 1993 study's own uncertainty is included the same comparison gives p = 0.063. So the replication was unlikely to confirm the printed gap, and its own estimate sits well below it. Both are true, and the page says both.

VI / The claimants' method, run on nothing

Does the 1993 procedure manufacture the claim?

The replication raised the possibility that the 1993 result was a false positive produced by the analysis:

Ericsson et al.'s [1] method of conducting two ANOVAs per question, and in particular comparing the best and good violinists but using the full sample degrees of freedom, increases the chances of finding p < 0.05.Macnamara and Maitra (2019), p. 17.

That can be tested. Feed the 1993 procedure data in which the three groups have the same true mean, and count how often it reports the claim anyway. What such a test can show turns on something the 1993 paper never printed: how spread out each of its groups was. A test that pools one error term across three groups is correctly sized when the groups share a spread, so null data built with one shared spread cannot show the worry at all. The first two families below are built that way; the last two keep each 2019 group's own spread, and the page starts on the first of those.

20,000 null datasets through the 1993 procedure

Null data

20,000 datasets in the 1993 design: each group of 10 drawn with replacement from its own group's 2019 residuals, so each keeps its 2019 spread.

best over good significant: 1,166 of 20,000 (5.83%)

complete correspondence (both contrasts significant, both in the claimed direction): 0 of 20,000 (0.00%)

The best-over-good contrast in every null dataset. Red: significant at .05 in the direction the 1993 paper claimed.

With one shared spread the procedure is correctly sized for best over good: it calls that step significant in 2.49% of null datasets with shuffled labels and 2.25% with the 1993 design, where a correctly sized test pointing one way should do so about 2.5% of the time. Keep each 2019 group's own spread and it is not: 5.83% with resampled residuals and 4.66% with normal draws, 1.9 to 2.3 times the nominal rate. The mechanism is the one the replication named. In 2019 the least accomplished group was the tightest (standard deviation 2,379 hours, against 3,355 and 5,348), so pooling it in shrinks the error term that judges best against good, and the step is then judged on the full sample's degrees of freedom.

Complete correspondence stays rare under every family (at most 0.11%), because the second contrast turns conservative when the tight group is the one it sets apart: alone it fires in 0.57% and 1.18% of the datasets that keep the 2019 spreads. The 1993 paper printed only the pooled spread, so which of these describes its own test cannot be settled from the record. All of this is this page's computation, on these two contrasts only, not the paper's other analyses.

VII / The second layer

What a 13-violinist replication could and could not see

1. The asymmetric instrument

Plant every gap from zero to 6,000 hours into the 2019 rows, resample, and count how often the replication's analysis confirms it (green). On the same axis, the dashed blue line is the p-value for the replication's actual estimate tested against each planted gap: where it falls below 0.05, the replication's data say the gap is not that large.

Bootstrap power against the planted best-over-good gap (green), and the p-value for the replication's own estimate against each gap (dashed blue). The orange marks are the three readings of the claimed size.

The curve is computed in your browser after the page loads.

The curve's left end is its calibration. With no gap planted it reads 0.035 (4,000 resampled replications at each of 25 gaps), against the 2.5% a correctly sized one-directional test would give: drawing 13 violinists with replacement from 13 makes the bootstrap a little generous, so the bootstrap powers on this page, if anything, overstate what the replication could see.

At the printed 2,109-hour gap the 2019 replication had about a one-in-four chance of confirming it, yet its estimate lies 2.1 standard errors below it.

The instrument is lopsided in a way neither paper says. It was unlikely to confirm a gap of the printed size, yet its own estimate rejects a gap that large (p = 0.044), though not once the 1993 study's uncertainty is included (p = 0.063). Every gap up to 1,994 hours lies inside the replication's own 95% interval for its estimate, so its data rule none of them out, and at that largest size it confirms the gap in only 23% of resampled replications. Between zero and 1,994 hours, the gaps the replication's own data cannot rule out, this replication can say almost nothing.

Operate it: any planted gap, any number of violinists

Planted 2,109 h into a copy of the 2019 rows; 13 best and 13 good violinists drawn from it with replacement, 10,000 times, through the replication's analysis: 2,489 confirm best over good (power 0.249, 25%).

The replication's own estimate (13 a group) against a true gap of 2,109 h: t(24) = −2.13, p = 0.044.

Groups larger than 13 are drawn with replacement from the 13 planted histories of each group, so the instrument assumes a larger study would find violinists like these. Group sizes run from 13 to 80; the gap from zero to 6,000 hours.

To give an 80% chance of confirming the gap, normal theory (two-sided .05, the page's computation) says a replication would need 72 violinists a group at the printed hours (d = 0.472 in the 2019 best-and-good spread) and 24 at the standardized reading (d = 0.836 in that spread); if the gap were d = 0.958 in the new study's own spread, as it was in 1993, 19 would do. Set the instrument to 72 a group at the printed gap and resampling gives 0.83, a little above the target, as the calibration above would predict.

2. Two studies, one gap

Units
Best minus good in each study, with 95% intervals: 1993 from its pooled standard deviation, 2019 from its rows (Student).

1993: 2,109 h, standard error 984 h

2019: −1,620 h, standard error 1,751 h

difference 3,729 h, z = 1.86, p = 0.063

fixed-effect pooled (the page's own): 1,213 h, standard error 858 h, p = 0.157

In hours the two studies differ at z = 1.86 (p = 0.063); in standardized units, z = 2.14 (p = 0.032). Pooled with fixed weights, the gap is 1,213 hours with a standard error of 858 (p = 0.157), or d = 0.18 with a standard error of 0.30 (p = 0.55). In standardized units each study uses the spread it has: 1993's d is the gap over the pooled spread of all three groups, the only spread it printed; 2019's is the gap over the pooled spread of its best and good groups. The pooling is this page's own, and it carries a caveat the arithmetic cannot see: these are different populations. The 1993 violinists were Berlin students with a mean age of 23.1 at interview; the 2019 violinists were Cleveland students recruited from 2014 to 2018, who reported more practice overall and far more competition entries.

3. The claimants' reply, operated

On the day the replication was published, The Guardian (21 August 2019) reported that Ericsson said the new paper actually replicated most of their findings. He said there were no objective differences between Macnamara’s best and good violinists, so no surprise they put in the same amount of practice. His co-author Ralf Krampe, the paper reported, said nothing in Macnamara’s paper made him question the original findings. He also told it: Do I believe that practice is everything and that the number of hours alone determine the level reached? No, I don’t, and But I still consider deliberate practice to be by far the most important factor. Ericsson then replied in print (Journal of Sports Sciences, 2020): the replication, he argued, did not reproduce the 1993 difference in objective performance between its best and good groups, and did not report or discuss that failure. The replication did print the comparison (successful competition entries, best against good, t = 1.55, p = 0.134, p. 11) and argued that its groups kept the same relative difference in skill (below); it did not call the comparison a failure to replicate. This page could not open the full text of that reply (the publisher returned HTTP 403), so it paraphrases it from the sentences indexed by Semantic Scholar's citation record, rather than quote it.

The 1993 paper supplies the yardstick itself: The best indicator of violin performance, besides the evaluation of the music professors, is success at open competitions. (p. 374). So put the two studies' competition records side by side.

Successful competition entries. 1993: the printed means for the best and good groups. 2019: recomputed from the released rows. Each panel has its own scale: the 2019 violinists entered far more competitions, so compare the gaps through the d values in the panel titles. They are close but not identical conventions: 1993's is the gap over the pooled spread of its contrast test (all three groups); 2019's is the paper's 2t/√df, which with the pooled spread of the best and good groups would be 0.61.

In 1993 the best violinists had 2.9 successful entries on average against 0.6 for the good, F(1, 27) = 19.35: by the paper's own contrast arithmetic a gap of d = 1.97. In 2019 the best had 13.31 against 8.69, t = 1.55, p = 0.134, d = 0.63. The 1993 best and good groups were far further apart on the skill marker, which is Ericsson's point. And the 2019 best group still led on competitions while trailing on practice, which is the replication's; its authors wrote: While our method of recruitment followed Ericsson et al. [1] and our violinists appeared to have the same relative difference in skill from each other, they may have an overall higher level of expertise than Ericsson et al.'s [1] violinists (e.g. the current violinists had entered many more competitions than those in 1993). The page does not fit a model that turns one gap into the other: two numbers, two readings, both printed. Macnamara's own summary: Practice makes you better than you were yesterday, most of the time, she told The Guardian. But it might not make you better than your neighbour. Or the other kid in your violin class.

4. The method on nothing

Section VI above. Complete correspondence stays rare under every null the page ran (at most 0.11%), so the 1993 procedure does not make the whole claim out of nothing. Its best-over-good step alone is another matter: correctly sized when the groups share one spread (2.49%), firing 1.9 to 2.3 times as often as a correctly sized test when they keep the 2019 spreads (5.83% and 4.66%). Which of these applied in 1993, the record cannot say.

VIII / The verdict, dated

OPEN

as of 2026-09-22; decided by: not yet; time to decision: not decided

What is open: whether a conservatory's best student violinists had accumulated more solitary practice by 18 than its good ones (the best-over-good step of the 1993 complete correspondence).

What came back: the step from the soloist-track students to the least accomplished. The replication confirmed it (Welch t = 3.26, p = 0.005), and the 1993 procedure run on its rows gives F = 11.45 against a printed 11.86.

Why the best-over-good step stays open: the replication was unlikely to confirm it at its printed size (planted, p = 0.240; power 0.25); the first author's printed objection, that the replication's best and good groups did not differ in skill as the 1993 groups did, is consistent with the replication's own competition records (d = 0.63 against 1.97), though its authors read those records differently, and no new data have been brought to it; and no second replication exists. Against the claim: the replication's estimate lies 2.13 standard errors below the printed gap (p = 0.044; 0.063 with the 1993 uncertainty included); and if the 1993 groups were spread as unequally as the 2019 groups, the 1993 test of this step would fire 1.9 to 2.3 times its nominal rate on null data (section VI). The evidence leans against a gap as large as printed without settling whether there is one. OPEN is this page's reading of that record.

Sources:

  1. Brooke N. Macnamara and Megha Maitra (2019). The role of deliberate practice in expert performance: revisiting Ericsson, Krampe & Tesch-Römer (1993). Royal Society Open Science 6(8): 190327. doi:10.1098/rsos.190327. CC BY 4.0.
  2. K. Anders Ericsson (2020). Towards a science of the acquisition of expert performance in sports: Clarifying the differences between deliberate practice and other types of practice. Journal of Sports Sciences 38(2): 159 to 176 (online 12 November 2019). doi:10.1080/02640414.2019.1688618
  3. K. Anders Ericsson and Kyle W. Harwell (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance: why the original definition matters and recommendations for future research. Frontiers in Psychology 10: 2396. doi:10.3389/fpsyg.2019.02396. CC BY 4.0.

What would change it: a preregistered direct replication whose best and good groups are separated by an objective performance measure (blind panel ratings or competition results), large enough for this step (by normal theory for an 80% chance, 72 violinists a group at the printed hours in the 2019 spread, 24 at the standardized reading, 19 if the gap is d = 0.958 in the new study's own spread); or the release of the 1993 participant rows.

The check

The two data files have not been hashed in this browser yet.

The claim

  • Pooled mean square from F = 4.59: 4,845,186.27 (Ericsson 2014 printed 4,845,186.27); SD 2,201.2 h. Predicted second F 11.857 against a printed 11.86; the rounding box (F to within half a unit in its last place, each mean to within half an hour, all corners) gives 11.828 to 11.885. With the teachers' mean moved to 3,520, the prediction is 11.06, outside, so the check can fail. It tests whether the two printed F values fit one pooled variance; a common shift or scale of all three means leaves both F values as they were, so it cannot vouch for the hours themselves.
  • Implied effect sizes: d = 0.958; three-group eta squared 0.379; partial eta squared 0.145 (first contrast) and 0.305 (second). The 1993 study's own post hoc power for its d at 10 a group, on its contrast's 27 error degrees of freedom: 0.54.

The replication, printed against recomputed (age 18)

quantityprintedrecomputed here
practice alone, means8224, 9844, 45588,224, 9,844, 4,557.64
best interval[6400, 10 048]6,400 to 10,048
good interval[6937, 12 751]6,937 to 12,751
least interval[3264, 5851]3,264 to 5,851
best against goodt = −0.93, p = 0.364, d = −0.38Student t = −0.925, p = 0.364, d = −0.38
good against leastt = 3.26 on 16.57 df, p = 0.005, d = 1.33Welch t = 3.26 on 16.57 df, p = 0.0048, d = 1.33
three groupschi-squared 13.90, p = 0.001, eta squared 0.26H = 13.90, p = 0.00096, eta squared 0.259
teacher-designed, means6251, 6821, 27996,251, 6,821, 2,799
teacher-designed, intervals[4293, 8210], [4482, 9160], [1862, 3735]4,293 to 8,210; 4,482 to 9,160; 1,862 to 3,735
teacher-designed, best against goodt = −0.37, p = 0.717, d = −0.15t = −0.37, p = 0.717, d = −0.15
teacher-designed, good against leastt = 3.13 on 15.75 df, p = 0.007, d = 1.28t = 3.13 on 15.75 df, p = 0.007, d = 1.28
teacher-designed, three groupschi-squared 10.74, p = 0.005, eta squared 0.23H = 10.74, p = 0.005, eta squared 0.230
competitions, means13.31, 8.69, 3.2313.31, 8.69, 3.23
competitions, pairst = 1.55, p = 0.134, d = 0.63; t = 2.79, p = 0.010t = 1.55, p = 0.134, d = 0.63; t = 2.79, p = 0.010

The weekly estimates times 52, summed from age 2, reproduce the authors' own accumulated totals for all 39 violinists; the page uses the weekly columns as input and the authors' totals only as this check.

Three readings the page shows once, as arithmetic

  • The 2019 paper (p. 17) says in 2014, Ericsson [20] revealed that the 95% confidence interval around the mean accumulated hours of practice alone to age 18 for the best violinists in Ericsson et al. [1] was 2894–11 926 h, and adds: It is perhaps surprising that Ericsson et al. found significant group differences between the best and good violinists in accumulated practice alone to age 18. Ericsson's 2014 sentence gives the range as an interval for individual values: Ericsson et al. (1993) did report F-tests so anyone would be able to calculate the pooled within-standard deviation (spooled), which was 2201 h and the 95% confidence interval around the mean of individual estimates of accumulated solitary practice at age 18 (M = 7410) ranged from 2894 to 11,926 h for individual values in the top group. It is 7,410 plus or minus 2.052 times 2,201.2, which is 2,893.6 to 11,926.4. A 95% interval for the mean of 10 would be about 7,410 plus or minus 1,428 hours.
  • The 1993 paper (p. 374) printed p < .05 beside F(1, 27) = 4.07 for the minutes of music the best and good violinists could play from memory; the p-value for that F is 0.054. It is not part of the claim this page tests.
  • The 2019 paper (p. 16) says Ericsson et al.'s [1] comparison of practice alone between the best and good violinists combined as a single group and the less accomplished violinists explained 48% of the variance in performance. It does not give the derivation, and the page cannot reproduce 48% from anything printed. Compared like with like, the three-group eta squared implied by the 1993 numbers is 0.379, against 0.259 in 2019, and 0.379 lies inside the replication's own printed interval for its value, [0.03, 0.44]. The replication's authors add: To be clear, explaining 26% of performance variance is not an inconsequential amount.

Every free choice, and every uncertainty

  • The claimed size has three readings; the default, hours as printed, is the one least favourable to the replication. The plant moves every best violinist by the same number of hours, spread evenly over their weekly estimates from their first practising year to 18, so each keeps their distance from the group mean. The words follow bootstrap power: below one half the replication was unlikely to confirm the planted gap, from one half to 80% it could have, and from 80% it probably would have. The copy carrying exactly the claimed gap is the median replication, and at every reading it comes back significant exactly when power passes one half.
  • The null section starts on the family that keeps each 2019 group's own spread, the only kind that can show the replication's Type I argument; the two shared-spread families are one click away.
  • Monte Carlo: 10,000 resampled replications per reading (seed 20190821), 20,000 null datasets per family, four families (seed 1993); change the plant seed and the power moves in the third decimal. The power curve uses fewer draws per point and says so as it runs.
  • Levene's test is read with deviations from the group mean; that reading picks the same two tests the paper printed. A blank or "." in a weekly column counts as zero, which is what reproduces the authors' totals.
  • The two populations differ in time, place, age at interview and competition history; the pooled estimate in section VII is the page's own and inherits all of that.
  • No number on this page is digitised from a figure. Figure 9 of the 1993 paper is described, not measured.
  • The page could not open the full text of Ericsson (2020), which returned HTTP 403, or the Dryad copy of the workbook, which returned HTTP 401. The OSF copy was used, and its SHA-256 matches the value OSF itself reports. The replication's electronic supplement (figshare, CC BY 4.0) was read on 2026-09-23: it holds no power analysis and no derivation of the 48%, and it is where the replication says it ran no test on its professionals.

Provenance

  • data/violinists-2019.csv: the 39 student rows of the replication's workbook (OSF project 4595q, file kqxgj, 382,719 bytes, SHA-256 970f7a56d2904044bd1234c6b1b837d520ee6d0c1a04a2ff5c75f2d66fc60dcf), reduced to group, age at interview, competitions, the authors' two totals, and the weekly estimates. No names, genders or other fields. SHA-256 4e11f9c29445a794727bd5b646373a69dccf65282f4a76a415ef7f7db2f306cc.
  • data/claim-1993.json: every number the 1993 paper printed about accumulated practice at 18 and the two skill markers. SHA-256 47d956db6371fe1ac0a8902c6e8e7d7c26bb89fc5c98c9abf5e8ffd86d5a96ee.
  • Licences: the same dataset was deposited by its authors at Dryad under CC0 1.0 (source 4 below); the OSF copy carries no licence field. The 2019 article is CC BY 4.0, as is Ericsson and Harwell (2019). The 1993 and 2014 articles and the newspaper are quoted briefly, with citation.

We searched Semantic Scholar’s record of the 99 works citing Macnamara and Maitra (2019), with their abstracts and the sentences that cite it, Europe PMC’s full-text index, Crossref’s bibliographic search, and the titles of OSF projects and preprints on 2026-09-22 and did not find a published power analysis of the replication’s best-versus-good comparison at the 1993 effect size, or a reanalysis that runs the 1993 orthogonal-contrast procedure on the released 2019 rows.

Sources

  1. K. Anders Ericsson, Ralf Th. Krampe and Clemens Tesch-Römer (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review 100(3): 363 to 406. doi:10.1037/0033-295X.100.3.363
  2. Brooke N. Macnamara and Megha Maitra (2019). The role of deliberate practice in expert performance: revisiting Ericsson, Krampe & Tesch-Römer (1993). Royal Society Open Science 6(8): 190327. doi:10.1098/rsos.190327. CC BY 4.0.
  3. Brooke N. Macnamara and Megha Maitra (2019). Additional Methods and Results from The role of deliberate practice in expert performance: revisiting Ericsson, Krampe & Tesch-Römer (1993). The Royal Society, electronic supplementary material to Royal Society Open Science 6: 190327. doi:10.6084/m9.figshare.9250445.v1. CC BY 4.0.
  4. Macnamara and Maitra, Replication and Extension: Seminal DP Study, OSF project 4595q, file kqxgj (the workbook); the same dataset deposited at Dryad, doi:10.5061/dryad.68db279, under CC0 1.0.
  5. Macnamara and Maitra, preregistration, OSF registration khjs7, registered 19 April 2017.
  6. K. Anders Ericsson (2020). Towards a science of the acquisition of expert performance in sports: Clarifying the differences between deliberate practice and other types of practice. Journal of Sports Sciences 38(2): 159 to 176 (online 12 November 2019). doi:10.1080/02640414.2019.1688618
  7. K. Anders Ericsson and Kyle W. Harwell (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance: why the original definition matters and recommendations for future research. Frontiers in Psychology 10: 2396. doi:10.3389/fpsyg.2019.02396. CC BY 4.0.
  8. K. Anders Ericsson (2014). Why expert performance is special and cannot be extrapolated from studies of performance in the general population: A response to criticisms. Intelligence 45: 81 to 103. doi:10.1016/j.intell.2013.12.001 (quoted from the in-press version)
  9. Ian Sample (2019). Blow to 10,000-hour rule as study finds practice doesn’t always make perfect. The Guardian, 21 August 2019.
  10. Florida State University News, 24 June 2020, notice of the death of K. Anders Ericsson (archived copy of 25 June 2020).