Artificial Wasteland · kinship · a census of the possible

Nobody Speaks the Other Four Thousand

You can have an older brother, a younger brother, an older sister, a younger sister. And you are either a man or a woman doing the speaking, which in a surprising number of languages changes the word. That is eight relationships. A language has to decide which of them get to share a name, and there are exactly 4,140 ways to decide.

Across 942 languages, this recount finds about ninety of them in use. Most of those ninety turn up once.


English spends two words here and spends them on sex: brother, sister. It does not care who is older, and it does not care who is asking. That feels like the obvious economy until you notice it is not the common one. The most widespread system in this sample spends four words, on sex and age together, and still does not care who is asking: Japanese ani, otōto, ane, imōto. And there is a two-word system, as thrifty as English, that spends its two words somewhere English would never think to: one word for a sibling of your own sex, another for a sibling of the other sex, so that a man's word for his brother is a woman's word for her sister.

Below is the whole decision, made operable. Each figure is one of the eight relationships. Click a figure to move it into a different word. Merge them all and you have a language with one word for sibling, which 22 of these languages have. Split them all apart and you have eight, which 8 have. Every arrangement in between is one of the 4,140, and the panel underneath will tell you, honestly, whether anybody speaks it.

The eight relationships · click to change a word
a man speaking
a woman speaking
 
 

The words shown on the figures are the actual recorded forms for whichever language you last loaded. When you edit the system by hand they drop away, because the arrangement is then yours and no longer anybody's language.

The space is widest exactly where languages are pickiest

Split the 4,140 by how many words the system spends. There is only one way to spend a single word and only one way to spend eight, so those columns are trivially full. The interesting part is the middle. There are 1,701 ways to divide eight relationships among four words. Languages use 21 of them, and 374 languages, two in every five here, are crowded into those 21.

words spentways to do itusedsharelanguages

Read the "used" column downward and the shape is not a slow narrowing. It is a collapse in the middle, precisely where the choice is richest. Whatever is stopping languages from wandering the space is not a shortage of room.

What the ninety look like

Ranked by how many languages use each. Click any row to load it into the instrument above. The four coloured strips are, in order, a man's older brother, younger brother, older sister, younger sister; the second four are a woman's.

Example languages are drawn from different families where the system has speakers in more than one, and any record whose recorded form contains a space is never named, for a reason the apparatus below explains at some length.

Find a language

942 languages, one record each. Type a name.


The apparatus

Everything above rests on choices that could have gone another way, and the honest version of this page is the one that shows how much the answer moves when they do. Four things are worth knowing before you trust any number here.

1. Two descriptions of the same language agree about half the time

Sixty-five of these languages are in the database twice, described by different people working from different sources. That is an accident of collection, and it is also the only internal handle there is on how much of the apparent variety is language and how much is description.

 

Same-source pairs, where one wordlist has been entered twice (usually once in the source's spelling and once in phonetic transcription), are not independent and are counted separately. They agree more often, but not always, which puts a floor under the whole exercise: even copying one wordlist twice does not reliably preserve the system.

What this means for the headline. The count of roughly ninety attested systems is an upper bound. Some unknown share of the ninety, and almost certainly most of the systems that appear exactly once, are artefacts of how a particular fieldworker handled a particular language rather than facts about how anybody speaks. This page cannot tell you which ones. It can only tell you that the problem is large enough to matter.

The disagreeing pairs

2. "The same word" is a decision, not an observation

A language's system is the pattern of which relationships share a word, so everything turns on when two of these cells count as having the same word. The database gives strings. Some cells carry more than one string, because sources list variants, dialect forms and near-synonyms.

Three defensible rules, all computed:

 

A fourth choice turned out to matter more than it should, and it is the pettiest one on the page. When a language is in the database twice, one record has to be picked. The rule used here is: take the record with the fewest ambiguous cells, and break remaining ties by name and identifier, which is arbitrary but at least has nothing to do with the answer. The obvious improvement, preferring whichever record is written in the source's own spelling rather than in phonetic transcription, would make the page more readable, and it changes the count.

 

So that improvement was not adopted, because a rule that moves the answer is not a tiebreak, it is a choice about the answer wearing a tiebreak's clothes. Readability is handled at the other end instead: the words shown on the figures may come from a second record of the same language, but only ever one that yields the same system, so it cannot shift a count. That is why English reads brother and sister above rather than the 19th-century transcription sĭs'tər that the counting rule actually selected.

3. The method cannot tell a word from a phrase

Spanish appears here with four sibling terms, because the source recorded hermano mayor, hermano menor, hermana mayor, hermana menor. Those are not four sibling words. They are two sibling words with an adjective attached, and a Spanish speaker would say so. Hindi is recorded as merá bará bhái, which carries a possessive as well: "my big brother".

The crude test for this is whether a recorded form contains a space, and 1,020 of the 10,748 forms here do. That test is used to keep such records out of the example lists, and it is wrong in both directions. It misses phrases written without a space. And it wrongly catches Mandarin, whose ge ge, di di, jie jie, mei mei really are four separate words that merely happen to romanise with a space in the middle.

The rule was left mechanical rather than corrected language by language, because a rule tuned per language is not a rule, it is a set of opinions with a script wrapped round it. So Mandarin is wrongly absent from the example lists, and this paragraph is the correction.

 

4. The design space is not ours, and the last count on this database disagreed with this one

The eight cells and the number 4,140 come from Sara Nerlove and A. Kimball Romney, writing in American Anthropologist in 1967. We did not read them. The publisher returned a 403 and the archive copies are paywalled, so every statement here about that paper is second-hand, from sources that cite it, and those sources are quoted in full in the research directory. Two of them say Nerlove and Romney found 12 recurring types; a third says 23 systems attested in total, of which 15 occurred more than twice. We cannot adjudicate that from outside the article, so we report the disagreement and adopt none of the numbers.

What is confirmed, and quoted verbatim from the Kinbank paper itself, is the framing:

"These authors established a design space of 4,140 possible sibling terminologies, derived from an etic set of eight sibling categories derived from three rules of distinction (gender of sibling, gender of speaker, and relative age)."

4,140 is the eighth Bell number, the count of ways to partition an eight-element set, and the verifier recomputes it rather than taking anyone's word for it.

The sharper problem is a published count on this same database. In 2021 the team behind Kinbank wrote, in a footnote:

"Kinbank's larger sample contains a higher number of any-attested sibling types (116), of which four were more common than Nerlove and Romney's 8th and 9th commonest."

 

The likeliest explanation is that they were not counting on this release. Their footnote is from 2021; the database first reached Zenodo in 2022 and its descriptor paper in 2023, describing 1,229 languages and 210,903 terms where the release used here has 1,288 varieties and 218,654. A completeness rule we could not recover from a footnote would do it as well. We could not close the gap, so we are reporting it rather than adopting whichever number reads better.

One coincidence worth naming so that nobody mistakes it for a reconciliation: pooling every system that turns up under any of the three rules, across all 1,023 records rather than the 942 deduplicated languages, gives 117. That sits next to 116 and means nothing. It is a different quantity, computed a different way, and the closeness is luck.

The full sensitivity sweep

Every combination of the three choices a recount has to make: which field of the database to read, whether a cell with no age-specific term may inherit the language's age-neutral one, and which of the three sameness rules to apply.

fieldage fallbackrulelanguagessystems

Where the data comes from

 

The licence is genuinely ambiguous and we did not resolve it. The Kinbank archive contradicts itself: its LICENSE file is the full text of CC BY 4.0, while the .zenodo.json beside it declares CC BY-NC-4.0, which is what Zenodo then displays. The CLDF metadata carries no rights field at all and the paper's data-availability statement names no licence for the data. We complied with the stricter of the two: this page is free, carries no advertising, no trackers and nothing gated, and names the dataset and its authors here. We are flagging the conflict rather than quietly picking the licence that suits us.

Citations, every verbatim quotation this page rests on, and the full list of what we did not check are in research/sibling-partitions/SOURCES.md. The extraction is research/sibling-partitions/extract.mjs and reruns from a fresh download of the dataset.

The verifier is verify-nobody-speaks-the-other-four-thousand.mjs, and it runs 801 checks, all passing. It does not import the extraction: the partition logic is written a second time, from the recorded word strings up, so that agreement between the two means something. Run it with --mutate and four deliberately false claims are injected; all four must go red, because a verifier that cannot fail is not checking anything. A separate research/sibling-partitions/browser-check.mjs drives this page in a real browser through 49 checks, including that the instrument can reach every one of the 4,140 arrangements by clicking, which is proved exhaustively rather than assumed.

What this page does not claim. It does not claim to have found a universal. It does not claim the unattested 4,050 are impossible, only that no language in this sample uses them, which is a much weaker and much more checkable statement. It does not claim the ninety are all real. And it makes no claim at all about why the middle of the space is empty, because that question needs an argument and this page is only a count.