The Prosody Workshop · a graph over the mouth

One Sound Away

A word ladder changes one letter at a rung and every rung has to be a real word. Move the rule off the page and into the mouth, and it changes one phoneme at a rung instead. This page builds that graph over 58,347 English words in your browser and walks it: the shortest chain of real words from any word to any other. Most of the time there is no chain, and the interesting part is that it can show you why rather than tell you.

Climb it

SCOWL's own size bands, smallest first. Widening the vocabulary adds rungs to climb on, and adds words that need climbing to.

Loading the dictionary…

Nothing above is looked up. The page fetches the dictionary and the same engine the lab ran, builds the graph in a background thread, and searches it. There is no stored table of answers to drift out of step with the prose.

The rule, and the one decision inside it

Two pronunciations are joined when you can turn one into the other by changing, dropping or adding exactly one phoneme. That is the standard one-phoneme metric of the psycholinguistics literature, and it is also, exactly, Lewis Carroll's ladder rule with the alphabet swapped out.

The decision worth naming is that a node here is a sound, not a word. cite, sight and site are one point in this space, not three, because no ear can hear a difference between them; a graph that made them three nodes would report a triangle of edges nobody can pronounce. Words with two readings sit in two places instead of one: either is both /IY DH ER/ and /AY DH ER/, and the two have different neighbours. So a rung is a sound, and the words printed beside it are every admitted spelling of that sound. The spelling is a label on the rung. It is not the rung.

Stress marks are stripped, because CMUdict writes AH0, AH1 and AH2 for one vowel at three stress grades and its non-primary digit is documented noise. Keeping the digits would cut the graph on an artefact. It turns out to matter very little either way: of 135,164 distinct pronunciations in the dictionary, exactly 304 collapse onto another when the digits go.

Where this is a replication, and where it is not

This graph is not new, and the honest thing is to say so before showing you anything. Vitevitch built it for English in 2008 over 19,340 words and reported its shape; Alderete, Mann and Tupper rebuilt it in 2025 at four times the size and released the edge lists. The 2008 paper even prints a ladder, as a passing illustration:

For example, to get from the word cat to the word dog, one can traverse the links between the nodes corresponding to the words bat, bag, and bog.

That chain is four rungs. It is the default in the instrument above, which answers three: cat, caught, cog, dog, taking the vowel of caught rather than going round through bat. He was illustrating, not claiming a shortest, and his lexicon was not this one. But it is a fair sample of what changes when you actually search instead of hand-picking.

So the shape of the graph below is a replication over a different lexicon, and it is labelled as one, with their numbers printed beside ours. What is new is narrower and is marked new here where it appears: the diameter of this graph and the pair that achieves it, a measurement of how much of the graph is the dictionary disagreeing with itself, a test of why the published isolate counts disagree, and the pathfinding instrument at the top, which we could not find anywhere else.

Most of English is islands

At the vocabulary the instrument opens on, the graph holds · distinct pronunciations of · words, joined by · edges. Those fall into · separate pieces. One of them is a continent holding · pronunciations. The second largest holds ·.

Everything else is dust. · pronunciations (·) have no neighbour at all: no single sound you can change, drop or add takes you to another English word. Counted by word rather than by sound, because a word with two readings is only alone if both are alone, that is · words.

Length is most of the story, and the staircase is clean. Two-phoneme sounds have a mean of 26.7 neighbours and not one of them is alone. By six phonemes the mean is 2.6 and a seventh of them are alone. By thirteen it is over half.

Share of pronunciations with no neighbour, by length in phonemes, at SCOWL band 60 (45,909 words). Lengths with fewer than 50 members are omitted.

But length is not all of it, and the short ones are the interesting ones. even is four phonemes long, /IY V IH N/, sits in the commonest band SCOWL has, and is completely alone. There is no English word at /IY V IH N/ minus a sound, plus a sound, or with one sound swapped. The obvious candidates all miss by exactly one sound too many. oven is /AH V AH N/, which differs in both of its vowels. seven and heaven are /S EH V AH N/ and /HH EH V AH N/: a phoneme longer, and both vowels different besides. evens would be a single addition, except that the dictionary gives it /IY V AH N Z/, with a different vowel in the second syllable, so it lands two off as well. Type even into the instrument above and it will show you the whole of what it can reach, which is nothing.

So does oomph (/UW M F/, three phonemes), angst, absurd, okra, asthma, puma. Vitevitch's 2008 examples were spinach and obtuse, and both are still alone here.

Isolation is a property of a word and a vocabulary, and it barely moves

A word is only an orphan relative to the answer pool it is being asked about, so the obvious worry is that the isolates are an artefact of a small dictionary. They are not. Growing the vocabulary from 45,909 words to the whole of SCOWL, 58,347 words, adds 12,438 possible partners and rescues 545 of the 10,300 isolated sounds. The other 9,755 are alone in the wider vocabulary too.

The longest ladder is 32 rungs new here

Inside the continent, every pair of sounds is connected by definition, so the question becomes how far apart two words can be. The answer is computed exhaustively rather than sampled: a breadth-first search from every one of the 19,980 nodes in the continent, all 399 million ordered pairs, no estimate anywhere. The average pair is 7.82 rungs apart. The single most distant pair is 32, and the pair is miscue and recruited.

Read the ladder and you can watch what it is made of. It leaves through the mis- / dis- / di- / re- prefixes, crosses the middle at read, rid and ring, and climbs back out through -ing and -ed. Nothing in the construction knows that English has prefixes and suffixes. The corridor is theirs anyway. And the rung in the middle is not a coincidence of the drawing: /R IH NG/ is the centre of the whole graph, the sound from which nothing else in the continent is more than 16 rungs away, and the longest ladder in English goes straight through it.

The stable number

The diameter is 32 at SCOWL band 35 (32,322 words) and still 32 at band 95 (58,347 words). Across that range the vocabulary grows by 80 per cent and the continent by 71 per cent, and the longest shortest path moves by one rung, at one band, and comes back. Whatever sets it is not the size of the dictionary.

Then type those two words into the instrument, and it says 31

Both numbers are right, and the gap between them is the whole point of making a rung a sound. The diameter is a distance between two sounds: /M IH S K Y UW/ and /R IH K R UW T IH D/ are 32 rungs apart and nothing in the continent is further. But recruited is written twice in the dictionary, and its other reading /R IY K R UW T IH D/ sits one rung nearer. The instrument is asked about words, so it takes the shortest route to either reading and answers 31.

Dropping one of those two numbers would read more cleanly and would be less true, so both are here. It is also the reason the check at the foot of this page re-walks the word-level 31 rather than the sound-level 32: a live check that quietly compared two different quantities would agree with itself and prove nothing.

A second wrinkle in that 32. Three of its rungs are the word resigning at three different attested pronunciations, and two are recruited at two. Those are real moves in sound and the graph is right to make them, but on the page they look like a stutter. Rebuilt with only each word's first dictionary reading, the continent shrinks from 19,980 sounds to 17,460 and the isolates rise from 10,300 to 11,636, so the choice is not free. It is stated below rather than buried.

The same graph on the page is a different country new here

Carroll's ladders run on spellings, so the fair comparison is to build the letter graph over exactly the same 45,909 words under exactly the same rule, one letter changed, dropped or added. The letter graph has exactly one node per word, so to compare like with like everything here is counted by word. English turns out to be considerably better connected in the mouth than on the page: 42.0 per cent of words are in the continent in sound against 31.2 per cent on the page, and 22.0 per cent of words are alone in sound against 30.2 per cent alone in spelling. The letter graph carries 59,130 edges where the sound graph carries 97,446, though those run over 45,909 and 49,451 nodes respectively, so the edge counts are the looser comparison of the three.

The comparison people have made before is between the two layers. What is easy to do once both graphs exist, and what we could not find done, is the comparison between the two sets of paths. Of the 113,502 word pairs that are one phoneme apart, 18,730 are not merely far apart in spelling: there is no chain of one-letter edits between them at all. And of the 59,130 pairs one letter apart, 4,083 have no sound chain at all, while 535 are homophones, one letter apart and zero sounds apart.

The extremes on each side are a matched pair, and both are ordinary English:

The widest disagreements between the two graphs, band 60. Left: adjacent in sound, distant on the page. Right: adjacent on the page, distant in the mouth.
one phoneme apartletter rungsone letter apartsound rungs

prevented and preventing differ by a single sound at the very end, /D/ against /NG/. On the page they are three letters apart and there is no way to walk between them one letter at a time in under thirty steps. recruited and recruiter differ by a single letter and are twenty-five rungs apart in the mouth. English spelling and English sound are two different maps of the same territory, and they disagree about what is next to what.

What the edges are made of new here

Every same-length edge in this graph is a minimal pair, so the edge list is a census of which oppositions English actually spends its vocabulary on. Before reading it, one correction has to be made, because without it the table is wrong in an interesting way.

4.4 per cent of the graph is the dictionary disagreeing with itself

4,265 of the 97,446 edges join two nodes that share a word: they are not two English words a sound apart, they are CMUdict hedging on how to transcribe one word. That is small overall and it is not spread evenly. It is concentrated almost entirely in the reduced vowels, and it lands on exactly the two entries a reader would otherwise find most striking. Raw, the second densest contrast in English looks like AH against IH with 1,269 pairs. 892 of those, 70 per cent, are one word written twice. IH against IY loses 545 of its 828 the same way.

The densest phoneme oppositions, band 60, before and after removing edges whose two sides are readings of one word. The vowel pairs are the ones that move.
contrastkindall edgesdistinct wordssame word

With that removed, the table is almost entirely the inflectional suffix system. D against Z at the top, 1,633 pairs, is -ed against -s. Third is D against NG, which is -ed against -ing. And of the 30,605 edges that add or drop a sound rather than swap one, the four commonest phones to appear or vanish are Z (7,024), S (4,306), D (2,865) and T (2,291). Those are not four suffixes but two pairs: /Z/ and /S/ are the consonantal shapes the plural, the possessive and the third person all take, and /D/ and /T/ are two of the three the past tense takes. Then ER (2,141), which is the agent ending.

Why the published numbers disagree new here

Here is a problem worth taking seriously. Vitevitch's 2008 network found 53.1 per cent of its words alone. Alderete and colleagues in 2025 found 43.0 per cent of word forms alone, and 47.3 per cent of lemmas. This graph finds 22.0 per cent. Those are the same construction and the same edge rule, and they differ by more than a factor of two. Somebody's number is going to get quoted as a fact about English.

The contrast table above suggests the answer. If the densest edges in the graph are -s, -ed and -ing, then the inflected forms are not extra vocabulary sitting on top of the graph. They are the connective tissue. And a 1964 pocket dictionary lists headwords, while SCOWL's word lists carry the inflections.

That is testable, so here is the test. Strip every regularly inflected form out of the lexicon, keeping a word only when it is not another admitted word plus -s, -es, -ed, -d or -ing in both its spelling and its pronunciation. That removes 18,444 words. Then remove 18,444 words at random as a control, so that "a smaller lexicon has more orphans" cannot explain the result on its own.

Removing the regular inflections against removing the same number of words at random (seed 20260816). Both cuts leave 27,465 words.
lexiconwordsedgescontinentwords alone

Cutting the lexicon at random takes the share of words with no neighbour from 22.0 to 34.1 per cent. Cutting exactly the regular inflections takes it to 47.3 per cent, which is squarely in the range the published figures occupy. The inflections carry far more of the graph than their number alone can account for.

So the honest reading of "how much of English has no phonological neighbour" is that it is not a fact about English until you say whether cats is a word. It is a fact about English and a lexicon, jointly, and the single largest lever is the one nobody states.

The check

Every number on this page is recomputed by research/phoneme-ladders/verify-one-sound-away.mjs, offline, from the dictionary rather than from any stored result. It re-derives the graph with its own arithmetic, re-runs the diameter sweep, and then asserts the shipped bytes. Specifically: