Where One Word Ends
Speech has no gaps between words, so a string of phonemes does not say where the words ended: ice cream and I scream are one sound, and so are nitrate and night rate. Type any phrase and this page enumerates every other thing it could have been, exhaustively and exactly, computed live in your browser from the shipped dictionary. Then the census, over 13,337 sentences of public-domain prose: 89.90 per cent of 171,593 word boundaries are recoverable from the phoneme string and a 42,495-word dictionary alone, with no grammar, meaning or context, and yet only 14.41 per cent of the sentences have a single reading. The Artificial Wasteland replicates Harrington and Johnstone (1987) at thirty times their lexicon and finds their reduced forms did more work than forty thousand extra words; runs the attested oronym list through a General American dictionary and classifies every failure; and counts, exhaustively, the 7.34 per cent of two-word phrases that can be completely re-seamed.
Say something, and see what else it was
loading the dictionary…
There are no gaps between words when you speak. The silences you think you hear are put there by you, after the fact, and a microphone will not find them. What leaves the mouth is one unbroken stream, so a string of speech sounds arrives at the ear without its word boundaries attached, and the listener has to invent them.
Mostly this is invisible, because mostly there is only one sensible way to cut. Sometimes there is not. ice cream and I scream are the same six phonemes in the same order: AY S K R IY M. So are nitrate and night rate. The technical name for the boundary that is doing nothing is internal open juncture; the popular name for a phrase with two readings is an oronym, widely attributed to Gyles Brandreth's The Joy of Lex of 1980, and disliked by some scholars because onomastics had already taken the word for the names of hills.
The box above is the question asked exactly. Type anything. It looks your words up in a pronouncing dictionary, glues the pronunciations into one string, throws the spaces away, and then finds every sequence of dictionary words that could have produced that string. It is doing this in your browser, from the same 1.2 MB dictionary file and the same engine this study ran, so nothing on this page is a stored answer that could have drifted away from the prose beside it.
Two things are true at once
Run that over real prose and you get a pair of numbers that seem to contradict each other, and do not.
The first: individual word boundaries are nearly always safe. Take a boundary in a real sentence and ask whether every valid reading of the whole string puts a boundary there. At a listener's vocabulary of 42,495 words, 89.90% of the 171,593 word boundaries in 13,337 sentences of public-domain prose survive that test. What decides them is the phoneme string and a dictionary. No grammar, no meaning, no context, no idea what the sentence is about.
The second: whole sentences are almost never safe. Only 14.41% of those same sentences have exactly one reading, and that share collapses as sentences get longer: 45.9% of three-word sentences, 15.9% at ten words, 2.3% at twenty, and 0.0% at thirty. A sentence is only unambiguous if every one of its seams is, and the one-in-ten that is not multiplies.
So the answer to "can you hear where words end" is yes, nearly always, one at a time and no, essentially never, all at once. That is not a paradox. It is what a product of many nearly-certain things looks like.
The dial, and the number moving
None of those numbers is a fact about English until you say which words count as words. A listener who knows 3,961 words has less to be confused by than one who knows 58,347. So the vocabulary is a dial, taken from SCOWL's own editorial size bands, and every figure is reported at every setting. The same 13,337 sentences are measured at all of them.
| Listener's vocabulary | Boundaries forced | False seams / 100 phones | Sentences with one reading | Mean readings | Median |
|---|---|---|---|---|---|
| 3,961 | 98.07% | 2.00 | 47.66% | 7.031 | 3 |
| 10,707 | 96.31% | 3.39 | 34.80% | 13.081 | 4 |
| 32,322 | 92.38% | 6.00 | 20.64% | 98.733 | 16 |
| 42,495 | 89.90% | 8.32 | 14.41% | 459.104 | 36 |
| 48,617 | 87.97% | 9.43 | 12.27% | 1,556.609 | 63 |
| 58,347 | 84.12% | 14.10 | 6.44% | 83,461.22 | 648 |
Every row is the same text. Forced is the share of real word boundaries that every reading of the string agrees on. False seams counts the places that were not word boundaries but could be read as one, per hundred phonemes. One reading is the share of sentences whose phoneme string has exactly one segmentation. Mean readings is over the subset of sentences in the 20 to 33 phoneme window described below, and counts word sequences rather than cuts.
This has been measured before, and the paper is worth reading
In 1987 Jonathan Harrington and Anne Johnstone published the same experiment in Computer Speech and Language, as part of the Edinburgh continuous speech recogniser. Their lexicon was the 4,000 highest-frequency words of the American Heritage Word Frequency Book, minus the letters of the alphabet, in RP citation forms, plus 5,300 fast-speech reduced forms derived from them by phonological rule: 9,300 pronunciations in one discrimination tree. Their corpus was 115 utterances transcribed by hand by a phonetician, averaging 7.07 words and 26.56 phonemes. They report, for the plain phonemic input, a mean of 1,790.9 word-strings per utterance.
Their example, printed in the paper, is worth having in front of you: the utterance branches are removed until there is just one left came out with "just under 16 000 alternative word-strings".
The comparison here is not a reproduction and the page does not pretend it is. Different material (written prose against transcribed speech), different dialect (General American against RP), different lexicon (an editorial band against a frequency list), and, most importantly, no reduced forms: this study glues citation forms together and models none of the reduction that real talkers perform. What can be compared is the quantity. Over the sentences whose phoneme strings fall in a window around their 26.56-phoneme average, at each vocabulary setting, the mean number of word sequences is the last column of the table above.
Read down it and their 1,790.9 sits between a 48,617-word vocabulary (1,556.609) and a 58,347-word one (83,461.22). Their 9,300 entries produced about as much ambiguity as roughly 48,617 citation-form words do here. The reduced forms were doing more work than forty thousand extra words. Which is the sharpest way to say the caveat this whole page rests on: the measurement below is of an idealised transcription of writing, and real speech is worse.
Lexical stress, the same test
They also found that marking lexical stress cut the ambiguity, from 1,790.9 word-strings to 475.8, a factor of 3.76. That test runs here too, on the same sentences and the same lexicon, once with CMUdict's stress digits kept and once with them stripped: 7.031 against 4.559, a factor of 1.54. The direction replicates. The size does not, and there are two honest reasons why: this lexicon has no reduced forms for stress to distinguish, and CMUdict's non-primary stress digit is known noise, measured elsewhere in this project's own rhyme work, where willow carries OW2 against pillow's OW0 for the same syllable.
And the boundaries, which is a different question again
The following year the same first author, with Gordon Watson and Maggie Cooper, asked the boundary question at COLING-88: how many word boundaries can be found from phoneme sequence constraints alone, the fact that some phoneme runs never occur inside a word. Their answer, over 1,411 boundaries in 145 utterances, was 523 of them, 37.1%, rising to 45.7% once one- and two-phoneme words and legal word-edge pairs were added.
The 89.90% above is not that number improved. It is a different criterion, and a much stronger one. Theirs asks a local question of a window slid along the string, against constraints precompiled from a lexicon; this one asks whether the boundary survives in every complete parse of the whole sentence against the dictionary itself. Global consistency beats a local window, and it should. What the gap does say is how much of word division is carried by knowing the words rather than by the phonotactics, on material that flatters both: their utterances were speech and ours are not.
Sixty-two oronyms somebody else collected
Everything above is this engine agreeing with itself. The test that matters is the one it could fail, so here is a list nobody here chose: the whole oronym entry of the rec.puzzles FAQ, compiled by Chris Cole and Matthew Daly, 62 phrase pairs and 19 sentence pairs, taken verbatim in the order the source prints them. These were assembled by English speakers, for fun, with no dictionary in the loop. They are a claim about the language, not about CMUdict.
50 of the 62 phrase pairs (80.6%) come out as literally the same string of phonemes, allowing every pronunciation variant on both sides; 43 of them are surfaced by the instrument at the top of this page at its default vocabulary. 2 contain a word the dictionary does not have. And 10 do not match at all.
That last group is the interesting one, and it is not a list of mistakes. Sort the failures by the shape of the difference, mechanically, and almost every one is a process that connected speech performs and a citation-form dictionary does not record:
| Shape of the difference | Pairs | One of them |
|---|---|---|
| vowel-quality | 7 | new direction / nude erection |
| degemination | 6 | homemaker / hoe-maker |
| affricate-seam | 4 | catch it / cat shit |
| flap | 3 | bee feeder / beef eater |
| voicing | 3 | standards-based / standard-spaced |
| yod | 1 | mature / much your |
| other | 1 | biggest hurdle / biggest turtle |
| h-dropping | 1 | stuffy nose / stuff he knows |
| stop-deletion | 1 | mint spy / mince pie |
Classified by rule, one edit operation at a time, from a shortest edit script between the two transcriptions. Where several edit scripts are equally short the choice among them is the aligner's and not a claim about what a speaker did, which is why biggest hurdle against biggest turtle scores as it does rather than as the doubled T plus dropped H a phonologist would write.
So night rain and night train differ by one T that the dictionary writes twice and a mouth says once. catch it and cat shit turn on whether CH is one sound or two. bee feeder and beef eater are the American flap, where T and D between vowels stop being different. tulips and two lips differ only in which reduced vowel the dictionary chose to write.
The jokes are not wrong and the dictionary is not wrong. The gap between them is the thing: the wordplay lives in exactly the places where speech departs from its own citation forms, which is why a page like this can measure a ceiling and never the real number.
Every pair, with its verdict
Words that are secretly phrases
nitrate is night plus rate with nothing left over. Over the whole lexicon at the default vocabulary, 45.93% of words (19,520 of 42,495) have a pronunciation that is exactly the concatenation of two or more other words' pronunciations. At the widest vocabulary it is 62.75%.
This is a stricter relation than the one the psycholinguistic literature usually counts. McQueen, Cutler, Briscoe and Norris reported in 1995 that 84% of polysyllabic English words contain a shorter word embedded somewhere inside them, over a 26,000-word lexicon. Embedding allows a remainder; exact concatenation does not, so the 45.93% here is a proper subset of a larger phenomenon and not a competing measurement of it. (That 84% is quoted from secondary sources; the original paper could not be opened from here, and this page says so rather than implying a reading it did not do.)
The strict reseam
Having another reading is cheap. two suns and too sons is not wordplay, it is spelling. The relation worth counting is the strong one: a second reading that shares not one of the original's interior boundaries, so that every seam has moved. ice cream against I scream is that. Call it a reseam.
Two counts. Over every distinct two- and three-word phrase that actually occurs in the three corpora, 9,536 of 227,030 (4.20%) have one. And exhaustively, over every one of the 114,639,849 ordered two-word phrases you can build from the commonest 10,707 English words, 7.34% can be completely reseamed. Neither is a search or a sample.
The ones in the books
Every phrase below occurs in Moby-Dick, Pride and Prejudice or the complete Shakespeare, with its reseam computed and ranked by the fewest and commonest words. Nothing here was chosen by hand.
Two words that are one word
From the exhaustive sweep over the commonest 3,961 words, the phrases whose entire other reading is a single word. At most one entry per starting word, so that seven ways to reseam undergraduate s- do not crowd out everything else; that is a presentation rule and the counts above are of everything.
The sweep also turned up art official, which is artificial, and which the machine that wrote this page is choosing to take as a compliment.
Two words that are two other words
The same sweep, longest first, where both readings are phrases.
What this is not
It is not a measurement of speech, and the difference is not a technicality.
Every string measured here is written English turned into citation forms: one dictionary pronunciation per word, glued without reduction, without assimilation across the join, without a syllable deleted. Talkers do all of those, which is why the attested oronyms above fail in exactly the places they do. Real speech is more ambiguous than this, not less, so 89.90% is a ceiling on what the phonemes can tell you and not an estimate of what they do tell you.
And real speech restores something this measure throws away on purpose. Listeners are not working from a bare phoneme string: oronym pairs are not acoustically identical, and the durational cues that mark an open juncture have been measured since Ilse Lehiste's 1960 monograph. Gow and Gordon showed in 1995 that hearers do not activate lips when it arrives inside tulips. Davis, Marslen-Wilson and Gaskell put it plainly: "The ambiguity created by embedded words is therefore not as severe as predicted by models based on phonemic representations." This page is a study of phoneme strings. It is not a study of hearing, and the two must not be quietly swapped.
Three more choices that move numbers, all declared. Stress digits are stripped except in the stress comparison, so AH0, AH1 and AH2 are one phoneme. The text's own pronunciation is CMUdict's first listing for each word; the listener's hypothesis space uses all variants. And counts include the identity reading, as Harrington and Johnstone's do, so a sentence with "one reading" means the true one and no other.
The corpora are what a public-domain shelf offers, and it shows. 41.1% of Moby-Dick's sentences, 18.3% of Pride and Prejudice's and 47.5% of Shakespeare's are dropped for containing a word CMUdict cannot pronounce, almost all of them proper names (stubb, longbourn, vincentio). Shakespeare in particular is here as a deliberately anachronistic control: CMUdict is General American of the 1990s and Early Modern English pronunciation was not that. It is reported separately for exactly that reason, and its numbers are a fact about running those texts through this dictionary, not about how anyone ever spoke.
The check
Everything above is recomputed by research/where-one-word-ends/verify-where-one-word-ends.mjs, 7,548 checks, and by mutate.mjs, which corrupts the engine 13 ways and requires the verifier to go red for each.
The load-bearing one is a second implementation written to disagree. brute.mjs enumerates every segmentation of a string and reads the answers off the list: the count is the length of the list, a seam is usable if some segmentation cuts there and forced if every one does. It is exponential and useless above about twenty phonemes, which is why it is a control and not the engine. The verifier runs both over 468 strings drawn from the corpora and the dictionary and requires exact agreement on all four quantities, 7,380 assertions in total.
The page you are reading is checked separately from the lab that fed it, in a real browser at 3 viewports: 33 checks that the instrument computes what the lab computes, phrase by phrase, that the vocabulary dial changes the answer, that a word the dictionary lacks is refused rather than guessed, and that nothing overflows or throws (headless.mjs). It caught a real error in the prose above: the first draft said ice cream was seven phonemes, and the ribbon drew six.
The page's own engine is a byte-identical copy of the lab's (seams.mjs, SHA-256 11e3875ad2174fa2), and the dictionary it ships is a byte-identical copy of the one One Sound Away ships, so the two instruments provably stand on one lexicon. Both copies are asserted.
Sources are vendored and pinned: CMUdict (research/prosody-workshop/cmudict.dict, 135,166 lines, 126,052 headwords), SCOWL's size bands, and the three Gutenberg texts by SHA-256. The data files behind every figure are research/where-one-word-ends/data/.
What was already known, and by whom
An adversarial scout was sent to kill this page's claims before a line of it was written, and killed two outright. Both are named here rather than worked around.
- The instrument is not new. Jennifer G. Hughes built an exhaustive oronym generator over CMUdict for her 2013 Cal Poly master's thesis, Misheard Me Oronyminator, walking a phonetic dictionary tree to find every valid word sequence matching a phonemic string, exactly the algorithm here. Harrington and Johnstone had built the same tree-structured enumerator in 1987. A live browser version exists too: evashort's Homophone Generator, which is deliberately not exhaustive and not exact, and says so in its own README. What is claimed here is only the conjunction: exhaustive, exact, in the reader's browser, over a stated public lexicon, with the counts done in arbitrary precision so that "showing 200 of 4" is a true sentence.
- The census is a replication. Harrington and Johnstone (1987) published the distribution of readings per utterance. This is that experiment at roughly thirty times the lexicon, on different material, with the vocabulary swept rather than fixed. Their figures are quoted above.
- The boundary question is older still. Harrington, Watson and Cooper (1988), and behind them Lamel and Zue (1984). Their measure is local and lexicon-free; this one is global. The numbers are not directly comparable and are not presented as though they were.
- The embedding statistic is theirs. McQueen, Cutler, Briscoe and Norris (1995) on 84% of polysyllabic words containing an embedded word; Norris, McQueen, Cutler and Butterfield (1997) on the Possible-Word Constraint, whose whole purpose is suppressing the spurious tilings this page counts.
- The typology is described elsewhere. Siham Mohammed Hasan Alkawwaz, "A Phono-Rhetorical Study of Oronyms in English", Academic Journal of Interdisciplinary Studies 10(2), 2021, sets out a word-to-phrase against phrase-to-phrase distinction on twelve hand-picked examples. No exhaustive count of strict reseams was found; the searches that failed to find one are listed in the repository README, because an absence is only evidence if you say what you looked for.