The Artificial Wasteland · measured on 2,193,841 words
Two Words in Every Five
Everyone has a position on English spelling. Almost nobody has a number. Here are four, computed from the pronouncing dictionary and two million words of running text: what a phonemic spelling would destroy, what it would rescue, what it would save you in letters, and the reason it keeps not happening.
The proposal is old and it is always the same. English spelling is a museum of accidents, so stop writing the history and write the sound. Spell it as it sounds.
The argument against it is also always the same, and it is about homophones: if you write the sound, then to, too and two become one word on the page, and something is lost. The argument for it is about waste: all those silent letters, all that time.
Both sides are making quantitative claims and neither side does the arithmetic. So this page does it. Nothing below is an opinion about spelling. Every figure comes out of two files that were already in this repository before the question was asked: the CMU Pronouncing Dictionary, and 2.19 million words of out-of-copyright prose.
First, operate it
Below is real text, respelt one symbol per sound. The scheme is written out further down; it is a plain one, and it does not mark stress. Watch what that hides, on the stress rung below. Type your own if you like.
Respell it
Flip the second switch. Nothing you did changed the language; you changed which of two readings the dictionary offers is taken as the one to write. The words that move are the ones the dictionary declines to choose between, and there are a great many of them.
The ladder
Take every word of the corpus that the dictionary knows, which is of , and ask a single question: does the dictionary give this spelling more than one pronunciation? For of those words the answer is yes.
That number is useless on its own, because most of those words are not ambiguous in any sense a writer would recognise. The is listed twice because it has a strong form and a weak one. Which is listed twice because the dictionary records both the hw of whine and the pronunciation without it. So the next step is to name the variations, one predicate each, and take them off the pile. Thirteen rules account for . What survives all thirteen is , and that residue was then read by hand, word by word, and is published below in full.
The bottom rung is the one the argument is about: a spelling that hides two different words. Read and read. The wind and to wind. A live wire and to live here. It is of running text, about one word in .
of running words have more than one reading in the dictionary
survive thirteen named variations
are genuine homographs, by hand audit
The thirteen rules
Each is a predicate over an alignment of two pronunciations of the same
spelling, in research/spell-it-as-it-sounds/classify.mjs. A pair is
explained when there exists an alignment every one of whose edits is named. The
percentages overlap, because one spelling can touch several rules, so they do not sum
to the total.
| rule | what it covers | words | of text |
|---|
What the merge costs
Now the other direction, and the argument everyone actually has. Merge every homophone: how much is lost?
Measured on the words the dictionary gives exactly one pronunciation for, so that no assumption has to be made about which reading a token carries, of words land in a class the sound cannot separate. That sounds like a catastrophe. It is not, and the reason is worth seeing: the classes are wildly lopsided.
| words | sound | spellings, commonest first | bits |
|---|
I and eye and aye share a sound, and 43,219 of the 43,772 occurrences are I. A reader who knows only the sound has almost no doubt left. Weight every class by how often it actually turns up and the whole loss is per word.
For scale: with no context to help, learning which word a token is costs in this corpus. So the entire homophone objection, the whole to/too/two case, is one part in of what a word tells you. It is not nothing. It is not much.
What it saves
The other claim is thrift. Words in this corpus average letters and sounds, so one symbol per sound is shorter.
Except that English has sounds and the Latin alphabet has 26 letters. Hand the 26 letters to the 26 commonest sounds and spell the rest with two letters each, and the saving falls to , or letters a word. That is a floor, not an estimate: a scheme that has to keep its digraphs unambiguous needs more symbols than this, not fewer, and this page has not surveyed the published schemes to say where any of them actually lands.
The obstacle nobody argues about
Here is the thing the arithmetic turns up that neither side of the argument talks about. To spell English as it sounds you must first decide whose English.
The dictionary used here is a dictionary of one accent, General American, built by hand at Carnegie Mellon. It is not neutral and does not claim to be. And even inside that one accent it declines to choose a reading for of running words.
Now step outside it. The dictionary is rhotic throughout: it writes the R of car, bird and water, because the speakers it models say it. The speakers of most of England, of Wales, of Australia, of New Zealand and of South Africa do not. That single difference touches of running words, distinct spellings, and it is a difference the dictionary cannot even express, so it appears nowhere in the ladder above.
Switch the first control on the instrument off and read the passage again. Nothing in that respelling is wrong. It is simply not the same language written down.
That is the real arithmetic of spelling reform. The homophone objection costs a word. The saving is under ten per cent of your letters. And the thing standing in the way is not that English sounds are hard to write, but that the question which sounds has never had an answer that every group of readers would accept.
The residue, in full
The spellings that survived all
thirteen rules, with the label each was given by hand and why. This table is a
judgement, made once, and published so that it can be argued with: change a label in
research/spell-it-as-it-sounds/residue-audit.tsv, re-run
measure.mjs, and every number on this page moves with it. The
commonest are shown.
| words | spelling | readings | label | note |
|---|
The check
Does it move if the corpus moves? The ladder was recomputed on two further bodies of text that share nothing with the Victorian novels. Peter Norvig's unigram counts from the Google Web Trillion Word Corpus, a different century and a different register, give . This website's own prose, written this decade, gives . Against the corpus's .
Could the instrument have said something else? The homophone measurement was re-run on the corpus after respelling it phonemically. A phonemic text has no homophone ambiguity left to find, so the answer has to come back exactly zero, and it does: merged words. A check that cannot fail is not a check.
Where the audit could be wrong. The bottom rung is a hand judgement over 277 spellings and the table above is all of it. Two calls to watch: the noun-verb stress pairs (a RECord, to reCORD) sit on the stress rung rather than the homograph rung, because a spelling that does not write stress leaves them exactly as ambiguous as the present one does; and where the same word simply has two accepted readings (either, route, aunt) it is called variation, not homography, because no information is carried by the choice.
What is assumed. That the CMU dictionary is a fair record of one accent (it is hand-built, it is not a survey, and its coverage of variants is uneven, which makes 39.28% a property of that dictionary and not a measurement of English). That corpus frequency is a reasonable weight for how often a reader meets a word. And that a reader has no context, which is false: real reading resolves homophones from the sentence, so 0.0708 bits is a ceiling on what the spelling channel carries alone, not an estimate of what a reader would lose. The rhoticity figure is a property of the dictionary's phoneme strings and needs no accent model.
Reproduce it.
node research/spell-it-as-it-sounds/measure.mjs recomputes every number
on this page from the two committed inputs and writes
data/results.json;
node research/spell-it-as-it-sounds/verify-two-words-in-every-five.mjs
re-derives them independently and asserts that the page and the data agree.
Neither reaches the network.
The respelling scheme
Declared, not proposed. Vowels: IY ee · IH i · EY ay · EH e · AE a · AA ah · AO aw · OW oh · UH uu · UW oo · AH u · ER er · AY ai · AW ow · OY oy. Consonants take their usual letters, with ch, j, sh, zh, th, dh, ng for the seven that have none. It uses digraphs and so is longer than the 26-letter floor computed above; it is chosen to be readable, which that floor is not.