Everything Arrives at Once
An English word spreads its distinctions along a line of about five phonemes. An American Sign Language sign stacks sixteen coded channels into about one moment. Take the same 1,977 concepts, coded once in each language, rub out one channel, and see which of them stop being tellable apart. The registered prediction was that each ASL channel would have to carry three times as much. It carries 1.32 times as much. The difference between the two codes turned out to be somewhere else entirely, and it is very large.
Two codes have the same job. Each has to keep the words of a language far enough apart that a listener, or a watcher, can tell which one arrived.
English does it along a line. night is N AY T, three phonemes in a row, and each phoneme is a little bundle of settings: where the tongue is, whether the vocal folds are running, how the air gets out. The word occupies time, and time is cheap, so English can afford to be leisurely: the average word in this study runs to 4.89 phonemes.
American Sign Language mostly does it in a single moment. In the standard annotation used here, a citation-form sign is a bundle of sixteen simultaneous settings: which fingers are selected, how far they are bent, whether they are spread, where the thumb is, whether the hand touches the body, where on the body, what the other hand is doing, whether there is a path movement and whether it repeats. Almost all of it arrives together. The average sign in this study occupies 1.10 of the sequential slots the annotation allows for.
So one code spends time and the other spends room. The obvious question is what the exchange rate is, and the way to ask it is to start taking things away.
Rub a channel out
Below are the same 1,977 concepts twice: on the left their ASL codes, on the right the pronunciations of their English glosses. Uncheck a channel and it stops being part of the code. Anything that was only ever kept apart by that channel now collapses into its neighbour, and the list underneath shows you exactly which things just became the same thing.
Start by changing nothing. The two numbers are already different, and that is the first finding.
The eraser
1,977 concepts, each coded twice. The score is how many of them still have a code nothing else shares.
American Sign Language
ASL-LEX 2.0, 16 channels, up to 6 sequential slots
—
English
CMUdict pronunciation, 10 distinctive-feature channels
—
What is no longer tellable apart
The first finding is in the baseline
With nothing erased, the full sixteen-channel ASL code leaves 1,720 of the 1,977 concepts with a code of their own: 87.00%. The English pronunciations leave 1,954, or 98.84%. Counted by items rather than by classes, 22.76% of the signs share their code with at least one other sign, against 2.33% of the words. It is roughly a factor of ten.
On the English side those collisions are the ordinary homophones, and they are real: night and knight, sun and son, wait and weight. Twenty-three groups, forty-six words.
On the ASL side there are 193 groups covering 450 signs, and the largest of them is street = way = plan = hall = bring = wide. Those are six different signs. Any signer will tell you so at a glance, and they would be right. They are the same row in this code, and that is precisely what is being measured. Hold on to that sentence, because it is the difference between what this page found and what it would be dishonest to claim it found. There is a section about it below.
Now take one channel away, and almost nothing happens
The classical way to price a contrast is functional load: merge it away and ask how much of the lexicon's entropy went with it. The measure goes back to Hockett in 1955 and was given its modern entropy form by Surendran and Niyogi in 2003. It has been computed for many spoken languages. As far as a literature search on the day this was written could establish, it had not been computed for a signed one.
The registered prediction was that ASL's heaviest channel would carry at least three times the load of English's heaviest, because simultaneity leaves the contrast nowhere else to go. It does not. ASL's heaviest channel is path movement at 0.0206. English's is consonant place at 0.0156. The ratio is 1.32.
Both numbers are small, and they are small for the same boring reason: a code with sixteen channels, or ten, can absorb the loss of any one of them. Delete path movement and 240 new groups form, but 1,573 of the 1,720 ASL classes survive intact. Marginal questions get marginal answers. The interesting question turned out to be the opposite one.
Keep exactly one channel, and everything happens
Everything from here to the exchange rate was not registered.
Five of the eight predictions broke, and what follows is what turned up afterwards, by
looking. It describes this dataset. It is not a tested hypothesis, and the right way to
treat it is as something for somebody else to test on another pair of languages. The one
piece of it that was registered in advance is named where it appears. The full accounting,
line by line, is in research/sign-lexicon-channels/DEVIATIONS.md.
Press solo on any row of the instrument and every other channel in that column is erased at once. This asks what a single channel can do on its own, which is the dual of asking what it costs to lose.
| code | channel | alone, it separates | and losing it costs |
|---|---|---|---|
| English | consonant place | 52.91% | 0.0156 |
| English | consonant manner | 52.76% | 0.0078 |
| English | vowel height | 31.56% | 0.0048 |
| ASL | nondominant handshape | 5.77% | 0.0003 |
| ASL | minor location | 5.46% | 0.0092 |
| ASL | second minor location | 3.03% | 0.0085 |
One channel of English, consonant place, tells more than half of these 1,977 concepts apart by itself. The best single channel of ASL manages one in seventeen. That gap is about nine to one, and it is not a gap in how much information the two codes carry overall, because with everything switched on the two are 87.00% and 98.84%. It is a gap in how the information is distributed.
Part of that 52.91% has to be handed back. An absent slot is itself a value: reading consonant place across a word also tells you how long the word is and which of its slots are vowels. That skeleton on its own separates 9.41%, and word length alone 0.66%. So consonant place is doing roughly forty-three points of real work on top of the skeleton, which does not change the conclusion but does change the number, and the number should be honest.
The mechanism is the slot budget, and it is measurable rather than rhetorical. English gets about 4.89 chances to say something with any one channel. ASL gets about 1.10. English does not need a channel to be rich, because it repeats. ASL cannot repeat, so it stacks.
The exchange rate
If English buys its distinctions with time, you can take the time away and see what a word is worth without it. Drag the slider under the English column: it hands the code only the first K phonemes of each word and throws the rest away.
One phoneme leaves 2.58% of the concepts distinguishable. Two leaves 23.93%. Three leaves 69.20%. Four leaves 91.96%. The whole simultaneous ASL sign sits at 87.00%, between the third phoneme and the fourth.
So the trade is legible: a whole ASL sign, sixteen channels arriving together in one moment, does about as much work as the first three and a bit phonemes of an English word. And it does it in 1.10 slots rather than 3.5. Simultaneity is buying a real compression in sequence, at the price of needing far more channels to be read at once.
This next paragraph is the registered half. The control asked the same thing from the other side: cut English to two phonemes and its heaviest channel should outweigh ASL's heaviest. It does, and not narrowly. At two phonemes English's consonant place carries 0.1221 against ASL's 0.0206, and at one phoneme it carries 0.2453. A one-slot English is far worse off than a one-slot ASL, which is exactly what sixteen simultaneous channels are for.
Three channels of English are free. One channel of ASL is.
A channel is free when erasing it merges nothing at all: every pair it separated was separated by something else too. Three of English's ten channels are free here, and only one of ASL's sixteen.
- Lip rounding is free, and for a reason worth stating: rounding is exactly determined by vowel height and backness together, with zero exceptions across all 1,977 words. That is the textbook claim that English rounding is not independently contrastive, arriving here as a count rather than an assertion.
- Lexical stress is free, and this one is an accident of the sample rather than a fact about English. Stress does distinguish real English pairs, but this concept set happens to contain no such pair, so the channel has nothing to do. Said plainly because the alternative is to let a null result pass as a finding.
- Syllabicity is free for a reason belonging to the coding, not the language: a vowel already carries a blank in every consonant channel, so the coding announces it twice.
- Major location is ASL's only free channel, and it is one of the four classical parameters of sign phonology. Erase it and not a single pair of the 1,977 merges. The obvious explanation, that minor location simply determines it, was tested and is false: 36 pairs share a minor location and differ in major location. It is determined by the other fifteen channels jointly, not by any one of them.
Counted as a share of each code, English can shed 30% of its channels at no cost and ASL 6%. That points the same way as everything else here: English has slack, and ASL does not.
Where the two codes agree
Erase channels greedily, always taking the one that destroys the most, and ask how many it takes to knock distinctness below half. ASL needs 5 of its 16. English needs 3 of its 10. As fractions those are 31.3% and 30.0%: two codes built on opposite principles, with almost exactly the same tolerance for damage. Nothing in the design forced that, and one run on one lexicon is not enough to call it a law, but it is the kind of coincidence worth writing down where the next person can find it.
The check
Eight predictions were registered in
research/sign-lexicon-channels/PREREGISTRATION.md and committed to the
repository before analyse.mjs existed. Three held. Five broke, including the
one this study was built around.
| registered | observed | |
|---|---|---|
| broke | ASL's heaviest channel is a location channel | path movement, 0.0206 |
| broke | English's heaviest is a vowel channel | consonant place, 0.0156 |
| broke | ASL's heaviest is at least 3x English's | 1.32x |
| held | ASL's baseline distinctness is the lower | 87.00% against 98.84% |
| broke | channels to half-collapse is at least 3 fewer for ASL | 5 against 3, the wrong way round |
| held | English cut to 2 phonemes outweighs ASL | 0.1221 against 0.0206 |
| held | frequency weighting does not change the top channel | unchanged in both, under three weightings |
| broke | one-handed signs carry more per channel than two-handed | 0.0151 against 0.0270, the wrong way round |
The last one is worth a sentence, because breaking it is informative. Two-handed signs are the more crowded half of the lexicon: 1,229 of them, 84.13% distinct, against 748 one-handed signs at 91.71%. Battison's symmetry and dominance conditions restrict what the second hand is allowed to do, and the restriction shows up here as a bill.
Robustness. Recoding ASL with the four classical parameters plus contact instead of the sixteen articulatory channels moves the baseline from 87.00% to 86.80% and leaves movement on top. Keeping all 2,380 entries including sign variants rather than one per English word moves it to 86.09%. Neither changes any conclusion above.
Run it yourself. Both programs are published in full at /checks/:
node research/sign-lexicon-channels/fetch-data.mjs
node research/sign-lexicon-channels/analyse.mjs
curl -sL https://artwaste.land/strata/everything-arrives-at-once/ \
-o public/strata/everything-arrives-at-once/index.html
curl -sL https://artwaste.land/strata/everything-arrives-at-once/lexicon.json \
-o public/strata/everything-arrives-at-once/lexicon.json
node verify-everything-arrives-at-once.mjs
The first pulls the two datasets from their publishers, about six megabytes, and refuses
any file whose sha256 is not the one recorded on the day this was written. They are not
shipped with the check on purpose: a check run against our copy of a dataset can only show
that we agree with ourselves. The second rebuilds both lexicons from those bytes, recomputes
every number on this page through the same channels-core.mjs your browser just
ran, and fails if any of them disagree. Seventy-four checks. The middle commands exist
because three groups of them compare the recomputation against things a bare download does
not carry: the analysis output, this page, and the payload your browser was handed. Skip them and the check says so in words and
counts those groups as not run, rather than shrinking quietly to a green tick. It also
carries --selftest, which plants ten defects and requires each to turn the real
battery red, plus a negative control that is expected to stay green because a partition
measure genuinely cannot see it.
This recipe was run the way a stranger would run it, in an empty directory holding nothing
but the published bytes, before this page shipped. That test is what found the first
version's fault: the fetcher's digest pins lived in a .json beside it, which
/checks/ does not publish, so the reader's very first command died on its own
first line while every check inside the workspace stayed green.
What this is a measurement of
Every number here is a property of a lexicon under a coding, and both codings are conventions made by people for their own purposes.
ASL-LEX's sixteen channels do not code palm orientation as such, and they do not code nonmanual marking, the face and posture that are part of many signs. They record a citation form. That is why street and wide land in the same row: not because a signer could confuse them, but because the things that separate them are not in this code. So the honest reading of the 22.76% is this is how finely the standard annotation resolves its own lexicon, and it is a useful thing to know about the annotation.
What survives that caveat is the part about shape rather than resolution. Adding the missing channels to ASL would add more channels at roughly one slot each; it would not give a sign four more slots. So the finding that English concentrates its power in few channels read many times, and ASL spreads it across many channels read once, is not an artefact of the annotation being coarse. The 8-to-1 solo gap could narrow. The reason for it would not move.
The English side has its own conventions: CMUdict is General American and gives one
pronunciation, so dialect is erased, and the feature table is one ordinary decomposition of
the ARPAbet inventory among several possible ones. It is written out in
lib/english-features.mjs so you can disagree with a named choice rather than
guess at one. The verifier checks that the table is lossless: no two of the 39 phones share
a feature bundle.
And the sample is a sample. These 1,977 concepts are the entries of ASL-LEX whose gloss is a single word CMUdict knows, one per gloss. They are not a random sample of ASL, and their glosses are not a random sample of English. They are a set of meanings a research group chose to film.