The fork that kept the time

The 1860 phonautograph had no clock. A tuning fork traced beside the voice was its only speed reference, and the number First Sounds assumed for that fork's pitch turned the earliest recorded human voice from a young girl in 2008 into an adult man in 2009. Same audio; one number about a fork; measured here, live, from the released MP3s.

On 27 March 2008, the front page of the New York Times carried a recording made ten years before Edison's phonograph. The recording is called a phonautogram: a wavy line scratched by a vibrating stylus onto lamp‑blacked paper, made by Édouard‑Léon Scott de Martinville in Paris on 9 April 1860. Scott never intended it to be played back. First Sounds, a collaborative of audio historians, extracted a sound from it anyway. The world heard, breathily and slowly, a voice singing Au clair de la lune. Almost every listener heard a young girl.

Fifteen months later, on 28 May 2009, First Sounds unveiled the same recording again. It was now half as fast, and half as high. The voice was an adult man. Nothing on the paper had changed. What had changed was one number about a fork.

The instrument nobody expected to hear

Scott's phonautograph was built for the eye, not the ear. It sent a stylus, attached to a membrane, across a sheet of paper darkened with soot, so a sound in the air became a wavy line you could look at. There was no clock in it, and no speed regulator. The paper moved because a human hand turned a crank. Scott himself wrote that the hand acquired an "almost uniform" speed with practice; in fact, the speed wandered so badly that on any uncorrected phonautogram a sung melody is unrecognisable by ear.

Scott's fix was elegant, and it is the whole trick that made First Sounds' work possible. He put a second stylus on a tuning fork, and let that stylus trace its own line on the same sheet, next to the voice. If you know the fork's frequency, its trace is a ruler in time: every cycle of the fork is one known interval, so you can renormalise the voice trace against it. The renormalised trace can be played back at a speed the crank never quite held.

That works only if you know the frequency of the fork. For most of the 1860 sheets Scott labelled the fork as "500 vibrations simples par seconde": five hundred simple vibrations per second. In modern English "vibration" and "cycle" are used interchangeably, so the phrase reads as 500 Hz. And that is how First Sounds first read it. On Au clair de la lune, at 500 Hz, they heard the natural speed of a young girl. It became the first recording released, and the story ran everywhere.

The correction: one word, half an octave, a different person

Early in 2009 they extracted a third phonautogram with a fork trace, this one of a speaking voice. At 500 Hz it sounded exactly like a tape played twice as fast. At 250 Hz, half the assumed frequency, it sounded like a low male voice at natural speed. Two other sheets, run at 250 Hz, became a plausible low male voice singing.

The problem was in one word. In mid‑nineteenth century French acoustics, a vibration simple was a half‑cycle, not a full one. Two vibrations simples equal one vibration composée, which is one cycle in our sense. Scott's "500 simple vibrations per second" therefore meant 500 half‑cycles per second: 250 Hz. Every phonautogram First Sounds had time‑corrected against Scott's fork had been played twice as fast as Scott's own paper marks intended. The person on the recording was not a girl. It was almost certainly Scott himself.

Try it below. The slider is the whole ambiguity: it sets what you assume the fork's tone was. The rest is arithmetic on the released audio.

250 Hz

Reading loading…

The slider does one thing. It multiplies the recording's playback speed, and therefore its pitch, by assumed_hz ÷ 250. If the true fork was 500 Hz, First Sounds' correction (which assumed 250 Hz) stretched every cycle to twice its real duration, so playback runs at twice the intended rate and every measured pitch reads twice as high. Everything else on the page is fixed: the actual voice pitch was measured from the First Sounds MP3 by autocorrelation, frame by frame, no smoothing, no interpolation from vocal expectations. What the reader sees change is only what that same voice would be if the fork had actually been at the frequency they choose.

How the pitch was measured

The blue histogram inside the chart is the distribution of fundamental frequency across every voiced frame of Au clair de la lune at the release's own calibration (the corrected, 250 Hz reading). Frames are 60 ms wide with a 30 ms hop; each voiced frame's fundamental was found by autocorrelation over 80–450 Hz, then re‑centred by parabolic interpolation on the peak. Frames whose RMS was under 15 % of the recording's overall RMS were dropped as silent or clicks. That yields voiced frames. The median across them, at the release's own calibration, is . That is the number the chart moves.

The reader sees no fitting to a preferred answer, and no smoothing. The three coloured bands on the chart come from decades of published speech‑science reference ranges: adult male fundamental frequency clusters around 85–180 Hz for speech, adult female around 165–255 Hz, prepubescent child around 250–400 Hz. Those bands overlap on the boundary, and I have drawn them exactly as they overlap.

The fork itself, measured

The whole argument rests on the fork actually being 250 Hz. The evidence for that is not just linguistic. Scott made a second recording, in 1859, of a diapason he calibrated himself: 2,613 stylus cycles counted in six seconds of paper travel, giving 435.5 Hz. That happens to be within a fraction of a Hertz of the Diapason Normal (the French standard tuning pitch legislated in 1859), which the acoustician Rudolph Koenig measured at 435.45 Hz. So Scott's own count on his own paper, and the state pitch of France in the year of the recording, agree to two decimal places.

The recording of that diapason still exists. It is on First Sounds' site as an MP3. Ask two independent methods to measure the fork tone in that MP3, a Fourier peak on the whole file and an autocorrelation on the same file after a narrow bandpass, and out come 435.62 Hz and 435.75 Hz. Two audio measurements, one paper count, one legal standard, all four inside a 0.3 Hz spread:

The check  ·  four independent numbers for one fork

1859 Diapason Normal (Koenig, cited by Feaster)435.45 Hz
Scott's own paper count (2,613 cycles ÷ 6 s)435.5 Hz
Fourier peak on the 1859 diapason MP3 (this page) Hz
Autocorrelation on the same MP3 (this page) Hz

Each number was made by a different instrument, on a different medium, in a different century. They agree to 0.3 Hz. The 250 Hz timecode fork of the 1860 phonautograms is the same instrument's other prong.

Where the honest edges of this showing are

Three things I did not do, and want stated, because their absence is what keeps the piece honest.

I did not re‑extract sound from Scott's paper. The audio measured on this page is the audio First Sounds released, already time‑corrected against Scott's fork trace at 250 Hz. What the slider does is invert that correction: at 500 Hz the recording plays at twice its released rate, at 125 Hz at half. That is arithmetic, not a re‑reading of the marks.

I did not detect the 250 Hz fork tone in the released voice mixes. First Sounds' released files strip out the fork trace by construction: its job was to time‑correct a separate line, then be discarded. Finding the fork in the voice mix would have been a doubled proof; I looked, and it is not cleanly there. The four numbers above are the evidence for the fork's frequency, not five.

I did not adjudicate whether the singer is Scott. First Sounds' own working conclusion, once the speed was corrected, was that the singer is likely Scott himself (Édouard‑Léon Scott de Martinville, then 43). That attribution rests on non‑acoustic circumstantial evidence (the sheet is signed and Scott almost always sang his own tests) that I cannot verify from a WAV. What I can say is that the released, 250 Hz‑corrected audio has a median voice fundamental around , in the ordinary range of adult male speech. At the 2008 reading, doubled, it lands in the range of a child. That is what the slider makes visible.

What the fork is doing, at the widest scale

Historians who have written about this correction usually treat it as a footnote: a small speed error, quickly fixed. It is worth widening the frame, because the fork is doing something more general. A recording is not just a series of pressure values. It is those values plus a clock that says which one comes next. Without the clock, the pressure values name a shape but not a sound. The phonautograph is the extreme case: no clock at all, only the fork's ruler beside the trace. Everything you know about how the recorded voice sounded (its pitch, its speed, its sex, the age of the person it came from) is downstream of one number about that ruler.

That is a much older situation than 1860. A radiocarbon date is a count times a half‑life; get the half‑life wrong by half an octave and every date in the museum lands in the wrong century. A spectrum from a distant star is a wavelength times a rest frame; get the rest frame wrong and every element in the star becomes a different element. Every measurement leans on a reference it does not itself contain, and if the reference drifts, everything downstream of it drifts silently with it. What is unusual about the phonautograph is that the drift has a face. A voice you were sure was a young girl becomes an unremarkable middle‑aged man, and the only thing that moved was the interpretation of one word.

The audio, and the rest of the sources

« The recording stylus sometimes left the paper and sometimes moved backwards along the time axis, violating basic assumptions of the "virtual stylus" approach and, for that matter, of sound recording in general. For this reason, we supposed at first that many of Scott's phonautograms, particularly the earliest ones, might remain permanently mute. » · First Sounds, notes to the Scott releases
« Adjusting the timecode to 250 Hz gave us a natural‑sounding low male voice for items 45 and 46 and a plausible (though lugubrious) item 36. This made a great deal of sense in retrospect: despite Scott's own statements to the contrary, "500 simple vibrations per second" would clearly have meant 250 Hz in the terminology of mid‑nineteenth century acoustics. Since then, the eduction of additional phonautograms has confirmed 250 Hz beyond reasonable doubt as Scott's timecode frequency. » · Patrick Feaster, "Édouard‑Léon Scott: An Annotated Discography," ARSC Journal 41:1 (Spring 2010), 49–50

All First Sounds audio is released under a Creative Commons Attribution licence: the Scott phonautograms, and the 1859 diapason in particular, are attributable to First Sounds (D. Giovannoni, P. Feaster, R. Nowak, M. Wittje, E. Cornell). The discography is the single source for every historical claim above; every number attributed to this page is recomputed from the released MP3s by a small verifier research/the-fork-that-kept-the-time/verify.py in the site's repository, which downloads the audio, runs the same measurements, and writes the JSON this page reads.