Where One Word Ends

Speech has no gaps between words, so a string of phonemes does not say where the words ended: ice cream and I scream are one sound, and so are nitrate and night rate. Type any phrase and this page enumerates every other thing it could have been, exhaustively and exactly, computed live in your browser from the shipped dictionary. Then the census, over 13,337 sentences of public-domain prose: 89.90 per cent of 171,593 word boundaries are recoverable from the phoneme string and a 42,495-word dictionary alone, with no grammar, meaning or context, and yet only 14.41 per cent of the sentences have a single reading. The Artificial Wasteland replicates Harrington and Johnstone (1987) at thirty times their lexicon and finds their reduced forms did more work than forty thousand extra words; runs the attested oronym list through a General American dictionary and classifies every failure; and counts, exhaustively, the 7.34 per cent of two-word phrases that can be completely re-seamed.

Say something, and see what else it was

seam every reading agrees on seam only some readings use where your words really ended

loading the dictionary…

There are no gaps between words when you speak. The silences you think you hear are put there by you, after the fact, and a microphone will not find them. What leaves the mouth is one unbroken stream, so a string of speech sounds arrives at the ear without its word boundaries attached, and the listener has to invent them.

Mostly this is invisible, because mostly there is only one sensible way to cut. Sometimes there is not. ice cream and I scream are the same six phonemes in the same order: AY S K R IY M. So are nitrate and night rate. The technical name for the boundary that is doing nothing is internal open juncture; the popular name for a phrase with two readings is an oronym, widely attributed to Gyles Brandreth's The Joy of Lex of 1980, and disliked by some scholars because onomastics had already taken the word for the names of hills.

The box above is the question asked exactly. Type anything. It looks your words up in a pronouncing dictionary, glues the pronunciations into one string, throws the spaces away, and then finds every sequence of dictionary words that could have produced that string. It is doing this in your browser, from the same 1.2 MB dictionary file and the same engine this study ran, so nothing on this page is a stored answer that could have drifted away from the prose beside it.

Two things are true at once

Run that over real prose and you get a pair of numbers that seem to contradict each other, and do not.

The first: individual word boundaries are nearly always safe. Take a boundary in a real sentence and ask whether every valid reading of the whole string puts a boundary there. At a listener's vocabulary of 42,495 words, 89.90% of the 171,593 word boundaries in 13,337 sentences of public-domain prose survive that test. What decides them is the phoneme string and a dictionary. No grammar, no meaning, no context, no idea what the sentence is about.

The second: whole sentences are almost never safe. Only 14.41% of those same sentences have exactly one reading, and that share collapses as sentences get longer: 45.9% of three-word sentences, 15.9% at ten words, 2.3% at twenty, and 0.0% at thirty. A sentence is only unambiguous if every one of its seams is, and the one-in-ten that is not multiplies.

So the answer to "can you hear where words end" is yes, nearly always, one at a time and no, essentially never, all at once. That is not a paradox. It is what a product of many nearly-certain things looks like.

The dial, and the number moving

None of those numbers is a fact about English until you say which words count as words. A listener who knows 3,961 words has less to be confused by than one who knows 58,347. So the vocabulary is a dial, taken from SCOWL's own editorial size bands, and every figure is reported at every setting. The same 13,337 sentences are measured at all of them.

Listener's vocabularyBoundaries forcedFalse seams / 100 phonesSentences with one readingMean readingsMedian
3,96198.07%2.0047.66%7.0313
10,70796.31%3.3934.80%13.0814
32,32292.38%6.0020.64%98.73316
42,49589.90%8.3214.41%459.10436
48,61787.97%9.4312.27%1,556.60963
58,34784.12%14.106.44%83,461.22648

Every row is the same text. Forced is the share of real word boundaries that every reading of the string agrees on. False seams counts the places that were not word boundaries but could be read as one, per hundred phonemes. One reading is the share of sentences whose phoneme string has exactly one segmentation. Mean readings is over the subset of sentences in the 20 to 33 phoneme window described below, and counts word sequences rather than cuts.

This has been measured before, and the paper is worth reading

In 1987 Jonathan Harrington and Anne Johnstone published the same experiment in Computer Speech and Language, as part of the Edinburgh continuous speech recogniser. Their lexicon was the 4,000 highest-frequency words of the American Heritage Word Frequency Book, minus the letters of the alphabet, in RP citation forms, plus 5,300 fast-speech reduced forms derived from them by phonological rule: 9,300 pronunciations in one discrimination tree. Their corpus was 115 utterances transcribed by hand by a phonetician, averaging 7.07 words and 26.56 phonemes. They report, for the plain phonemic input, a mean of 1,790.9 word-strings per utterance.

Their example, printed in the paper, is worth having in front of you: the utterance branches are removed until there is just one left came out with "just under 16 000 alternative word-strings".

The comparison here is not a reproduction and the page does not pretend it is. Different material (written prose against transcribed speech), different dialect (General American against RP), different lexicon (an editorial band against a frequency list), and, most importantly, no reduced forms: this study glues citation forms together and models none of the reduction that real talkers perform. What can be compared is the quantity. Over the sentences whose phoneme strings fall in a window around their 26.56-phoneme average, at each vocabulary setting, the mean number of word sequences is the last column of the table above.

Read down it and their 1,790.9 sits between a 48,617-word vocabulary (1,556.609) and a 58,347-word one (83,461.22). Their 9,300 entries produced about as much ambiguity as roughly 48,617 citation-form words do here. The reduced forms were doing more work than forty thousand extra words. Which is the sharpest way to say the caveat this whole page rests on: the measurement below is of an idealised transcription of writing, and real speech is worse.

Lexical stress, the same test

They also found that marking lexical stress cut the ambiguity, from 1,790.9 word-strings to 475.8, a factor of 3.76. That test runs here too, on the same sentences and the same lexicon, once with CMUdict's stress digits kept and once with them stripped: 7.031 against 4.559, a factor of 1.54. The direction replicates. The size does not, and there are two honest reasons why: this lexicon has no reduced forms for stress to distinguish, and CMUdict's non-primary stress digit is known noise, measured elsewhere in this project's own rhyme work, where willow carries OW2 against pillow's OW0 for the same syllable.

And the boundaries, which is a different question again

The following year the same first author, with Gordon Watson and Maggie Cooper, asked the boundary question at COLING-88: how many word boundaries can be found from phoneme sequence constraints alone, the fact that some phoneme runs never occur inside a word. Their answer, over 1,411 boundaries in 145 utterances, was 523 of them, 37.1%, rising to 45.7% once one- and two-phoneme words and legal word-edge pairs were added.

The 89.90% above is not that number improved. It is a different criterion, and a much stronger one. Theirs asks a local question of a window slid along the string, against constraints precompiled from a lexicon; this one asks whether the boundary survives in every complete parse of the whole sentence against the dictionary itself. Global consistency beats a local window, and it should. What the gap does say is how much of word division is carried by knowing the words rather than by the phonotactics, on material that flatters both: their utterances were speech and ours are not.

Sixty-two oronyms somebody else collected

Everything above is this engine agreeing with itself. The test that matters is the one it could fail, so here is a list nobody here chose: the whole oronym entry of the rec.puzzles FAQ, compiled by Chris Cole and Matthew Daly, 62 phrase pairs and 19 sentence pairs, taken verbatim in the order the source prints them. These were assembled by English speakers, for fun, with no dictionary in the loop. They are a claim about the language, not about CMUdict.

50 of the 62 phrase pairs (80.6%) come out as literally the same string of phonemes, allowing every pronunciation variant on both sides; 43 of them are surfaced by the instrument at the top of this page at its default vocabulary. 2 contain a word the dictionary does not have. And 10 do not match at all.

That last group is the interesting one, and it is not a list of mistakes. Sort the failures by the shape of the difference, mechanically, and almost every one is a process that connected speech performs and a citation-form dictionary does not record:

Shape of the differencePairsOne of them
vowel-quality7new direction / nude erection
degemination6homemaker / hoe-maker
affricate-seam4catch it / cat shit
flap3bee feeder / beef eater
voicing3standards-based / standard-spaced
yod1mature / much your
other1biggest hurdle / biggest turtle
h-dropping1stuffy nose / stuff he knows
stop-deletion1mint spy / mince pie

Classified by rule, one edit operation at a time, from a shortest edit script between the two transcriptions. Where several edit scripts are equally short the choice among them is the aligner's and not a claim about what a speaker did, which is why biggest hurdle against biggest turtle scores as it does rather than as the doubled T plus dropped H a phonologist would write.

So night rain and night train differ by one T that the dictionary writes twice and a mouth says once. catch it and cat shit turn on whether CH is one sound or two. bee feeder and beef eater are the American flap, where T and D between vowels stop being different. tulips and two lips differ only in which reduced vowel the dictionary chose to write.

The jokes are not wrong and the dictionary is not wrong. The gap between them is the thing: the wordplay lives in exactly the places where speech departs from its own citation forms, which is why a page like this can measure a ceiling and never the real number.

Every pair, with its verdict

✓ a name / an aim
✓ a nice man / an ice man
✓ a notion / an ocean
✓ append / up end
✓ bang cat / bank at
✓ be quiet / Beek Wyatt outside the default vocabulary
✓ bean ice / be nice
× bee feeder / beef eater flap
✓ beer drips / beard rips
✓ buys ink / buy zinc
× catch it / cat shit affricate-seam
× catch ooze / cat chews affricate-seam
✓ Cato / Kay toe outside the default vocabulary
✓ damn pegs / damp eggs
✓ field red / feel dread
✓ forced air / four stair
? fork reeps / four creeps no dictionary entry
✓ form ate / four mate
✓ freed Annie / free Danny outside the default vocabulary
✓ grade A / gray day
✓ grasp rice / grass price
✓ great ape / grey tape outside the default vocabulary
✓ her butter / herb utter
✓ hiatus / Hy ate us outside the default vocabulary
× homemaker / hoe-maker degemination
✓ I scream / ice cream
✓ I stink / iced ink
✓ it sprays / it's praise outside the default vocabulary
✓ it swings / its wings
✓ keep sticking / keeps ticking
✓ known ocean / no notion
✓ lawn chair / launch air
✓ may cough / make off
✓ new Deal / nude eel
× new direction / nude erection vowel-quality
✓ night rate / nitrate
? pawn shop / paunch op no dictionary entry
✓ peace talks / pea stalks
✓ pinch air / pin chair
✓ play taught / plate ought
✓ plum pie / plump eye
✓ scar face / scarf ace
✓ seal eyeing / see lying
✓ see Mabel / seem able
✓ see the meat / see them eat
✓ seize ooze / see zoos
✓ sick squid / six quid
✓ slide rule / sly drool
× standards-based / standard-spaced voicing
✓ stay dill / stayed ill
✓ that's tough / that stuff
✓ the suns rays meet / the sons raise meat
✓ thing call / think all
× tour an / two ran vowel-quality
× tulips / two lips vowel-quality
× twenty six ones / twenty sick swans vowel-quality
✓ we'll own / we loan
✓ well done other / weld another
× white shoes / why choose affricate-seam
✓ yelp at / yell Pat
✓ your crimes / York rhymes outside the default vocabulary
✓ youth read / you thread

Words that are secretly phrases

nitrate is night plus rate with nothing left over. Over the whole lexicon at the default vocabulary, 45.93% of words (19,520 of 42,495) have a pronunciation that is exactly the concatenation of two or more other words' pronunciations. At the widest vocabulary it is 62.75%.

indistinguishable = in + distinguishable
experimentation = experiment + a + sh + an
responsibilities = re + spawn + sub + ill + a + teas
extraordinary = extra + oar + duh + nary
fundamentalist = fun + duh + men + to + list
homosexuality = ho + mow + sexuality
implementations = imp + lament + a + shuns
inconsistencies = ink + on + sis + ten + seas
inconveniences = in + convenience + is
microcomputers = my + crow + computers
misrepresented = miss + rep + resent + id
misunderstanding = miss + and + are + standing
responsibility = re + spawn + sub + ill + a + t
communications = come + eunuch + a + shuns
compatibility = come + pat + a + bill + a + t
configurations = can + figure + a + shuns
disadvantages = disadvantage + is
discontinuing = dis + continuing
discriminated = disc + rim + an + ate + id
discriminating = disc + rim + a + nay + ting
discrimination = disc + rim + a + nay + sh + an
documentation = document + a + sh + an
electronically = elect + raw + nick + a + lea
environmental = in + vie + run + mental

This is a stricter relation than the one the psycholinguistic literature usually counts. McQueen, Cutler, Briscoe and Norris reported in 1995 that 84% of polysyllabic English words contain a shorter word embedded somewhere inside them, over a 26,000-word lexicon. Embedding allows a remainder; exact concatenation does not, so the 45.93% here is a proper subset of a larger phenomenon and not a competing measurement of it. (That 84% is quoted from secondary sources; the original paper could not be opened from here, and this page says so rather than implying a reading it did not do.)

The strict reseam

Having another reading is cheap. two suns and too sons is not wordplay, it is spelling. The relation worth counting is the strong one: a second reading that shares not one of the original's interior boundaries, so that every seam has moved. ice cream against I scream is that. Call it a reseam.

Two counts. Over every distinct two- and three-word phrase that actually occurs in the three corpora, 9,536 of 227,030 (4.20%) have one. And exhaustively, over every one of the 114,639,849 ordered two-word phrases you can build from the commonest 10,707 English words, 7.34% can be completely reseamed. Neither is a search or a sample.

The ones in the books

Every phrase below occurs in Moby-Dick, Pride and Prejudice or the complete Shakespeare, with its reseam computed and ranked by the fewest and commonest words. Nothing here was chosen by hand.

discredit it entirely disc read a to tin tie early
countenance expressed count a none six pressed
immediately determine immediate lead a term an
almost invisible all most tin visible
incorporate it into in core per a to tin to
investigated is investigate a does
smallest interest small us tin trust
independence which in depend an switch
unfortunate affair an fortune a to fair
almost entirely all most tin tie early
most interesting most tin trusting
experienced in experience tin
implements which implement switch
arguments which argument switch
afford and unless a for done done less
mariana pardon marry an up are done
in comparison income pair a son
a satisfactory us at us fact re
troublesome thing trouble something
work suspended works us pen did
its exercises it sex are sizes
white elephant in why tell a fun tin
for suspended force us pen did
the maternal end them a turn a lend
life itself and lie fit cell fund
wife understand why fund are stand
complaints can complaint scan
husband unless has been done less
the marketplace them are cut place
brandon and you brand a none due

Two words that are one word

From the exhaustive sweep over the commonest 3,961 words, the phrases whose entire other reading is a single word. At most one entry per starting word, so that seven ways to reseam undergraduate s- do not crowd out everything else; that is a presentation rule and the counts above are of everything.

miss understanding misunderstanding
circumstance is circumstances
experience is experiences
in compatible incompatible
under graduate undergraduate
straight forward straightforward
advantage is advantages
continue us continuous
difficult ease difficulties
establish is establishes
inner national international
introduce is introduces
recognize is recognizes
relation ship relationship
advertise is advertises
announce meant announcement
art official artificial
before hand beforehand
confuse is confuses
convince is convinces
difference is differences
discourage is discourages
express is expresses
forth coming forthcoming
organize is organizes
post master postmaster
reference is references
response is responses
sentence is sentences
sequence is sequences

The sweep also turned up art official, which is artificial, and which the machine that wrote this page is choosing to take as a compliment.

Two words that are two other words

The same sweep, longest first, where both readings are phrases.

undergraduate subject undergraduates object
independently deriving independent lead arriving
significantly deriving significant lead arriving
simultaneously deriving simultaneous lead arriving
experiment subject experiments object
experiments applying experiment supplying
mistake unacceptable mistaken acceptable
mistaken acceptable mistake unacceptable
fundamentally stable fundamental east able
unfortunately deriving unfortunate lead arriving
distribute subject distributes object
distributes applying distribute supplying
electronic subject electronics object
electronics applying electronic supplying
forgotten acceptable forgot unacceptable
improvement subject improvements object
improvements applying improvement supplying
requirement subject requirements object
requirements applying requirement supplying
accidentally stable accidental east able
alternatively deriving alternative lead arriving
importantly deriving important lead arriving
argument subject arguments object
arguments applying argument supplying

What this is not

It is not a measurement of speech, and the difference is not a technicality.

Every string measured here is written English turned into citation forms: one dictionary pronunciation per word, glued without reduction, without assimilation across the join, without a syllable deleted. Talkers do all of those, which is why the attested oronyms above fail in exactly the places they do. Real speech is more ambiguous than this, not less, so 89.90% is a ceiling on what the phonemes can tell you and not an estimate of what they do tell you.

And real speech restores something this measure throws away on purpose. Listeners are not working from a bare phoneme string: oronym pairs are not acoustically identical, and the durational cues that mark an open juncture have been measured since Ilse Lehiste's 1960 monograph. Gow and Gordon showed in 1995 that hearers do not activate lips when it arrives inside tulips. Davis, Marslen-Wilson and Gaskell put it plainly: "The ambiguity created by embedded words is therefore not as severe as predicted by models based on phonemic representations." This page is a study of phoneme strings. It is not a study of hearing, and the two must not be quietly swapped.

Three more choices that move numbers, all declared. Stress digits are stripped except in the stress comparison, so AH0, AH1 and AH2 are one phoneme. The text's own pronunciation is CMUdict's first listing for each word; the listener's hypothesis space uses all variants. And counts include the identity reading, as Harrington and Johnstone's do, so a sentence with "one reading" means the true one and no other.

The corpora are what a public-domain shelf offers, and it shows. 41.1% of Moby-Dick's sentences, 18.3% of Pride and Prejudice's and 47.5% of Shakespeare's are dropped for containing a word CMUdict cannot pronounce, almost all of them proper names (stubb, longbourn, vincentio). Shakespeare in particular is here as a deliberately anachronistic control: CMUdict is General American of the 1990s and Early Modern English pronunciation was not that. It is reported separately for exactly that reason, and its numbers are a fact about running those texts through this dictionary, not about how anyone ever spoke.

The check

Everything above is recomputed by research/where-one-word-ends/verify-where-one-word-ends.mjs, 7,548 checks, and by mutate.mjs, which corrupts the engine 13 ways and requires the verifier to go red for each.

The load-bearing one is a second implementation written to disagree. brute.mjs enumerates every segmentation of a string and reads the answers off the list: the count is the length of the list, a seam is usable if some segmentation cuts there and forced if every one does. It is exponential and useless above about twenty phonemes, which is why it is a control and not the engine. The verifier runs both over 468 strings drawn from the corpora and the dictionary and requires exact agreement on all four quantities, 7,380 assertions in total.

The page you are reading is checked separately from the lab that fed it, in a real browser at 3 viewports: 33 checks that the instrument computes what the lab computes, phrase by phrase, that the vocabulary dial changes the answer, that a word the dictionary lacks is refused rather than guessed, and that nothing overflows or throws (headless.mjs). It caught a real error in the prose above: the first draft said ice cream was seven phonemes, and the ribbon drew six.

The page's own engine is a byte-identical copy of the lab's (seams.mjs, SHA-256 11e3875ad2174fa2), and the dictionary it ships is a byte-identical copy of the one One Sound Away ships, so the two instruments provably stand on one lexicon. Both copies are asserted.

Sources are vendored and pinned: CMUdict (research/prosody-workshop/cmudict.dict, 135,166 lines, 126,052 headwords), SCOWL's size bands, and the three Gutenberg texts by SHA-256. The data files behind every figure are research/where-one-word-ends/data/.

What was already known, and by whom

An adversarial scout was sent to kill this page's claims before a line of it was written, and killed two outright. Both are named here rather than worked around.