You can have an older brother, a younger brother, an older sister, a younger sister. And you are either a man or a woman doing the speaking, which in a surprising number of languages changes the word. That is eight relationships. A language has to decide which of them get to share a name, and there are exactly 4,140 ways to decide.
Across 942 languages, this recount finds about ninety of them in use. Most of those ninety turn up once.
English spends two words here and spends them on sex: brother, sister. It does not care who is older, and it does not care who is asking. That feels like the obvious economy until you notice it is not the common one. The most widespread system in this sample spends four words, on sex and age together, and still does not care who is asking: Japanese ani, otōto, ane, imōto. And there is a two-word system, as thrifty as English, that spends its two words somewhere English would never think to: one word for a sibling of your own sex, another for a sibling of the other sex, so that a man's word for his brother is a woman's word for her sister.
Below is the whole decision, made operable. Each figure is one of the eight relationships. Click a figure to move it into a different word. Merge them all and you have a language with one word for sibling, which 22 of these languages have. Split them all apart and you have eight, which 8 have. Every arrangement in between is one of the 4,140, and the panel underneath will tell you, honestly, whether anybody speaks it.
The words shown on the figures are the actual recorded forms for whichever language you last loaded. When you edit the system by hand they drop away, because the arrangement is then yours and no longer anybody's language.
Split the 4,140 by how many words the system spends. There is only one way to spend a single word and only one way to spend eight, so those columns are trivially full. The interesting part is the middle. There are 1,701 ways to divide eight relationships among four words. Languages use 21 of them, and 374 languages, two in every five here, are crowded into those 21.
| words spent | ways to do it | used | share | languages |
|---|
Read the "used" column downward and the shape is not a slow narrowing. It is a collapse in the middle, precisely where the choice is richest. Whatever is stopping languages from wandering the space is not a shortage of room.
Ranked by how many languages use each. Click any row to load it into the instrument above. The four coloured strips are, in order, a man's older brother, younger brother, older sister, younger sister; the second four are a woman's.
Example languages are drawn from different families where the system has speakers in more than one, and any record whose recorded form contains a space is never named, for a reason the apparatus below explains at some length.
942 languages, one record each. Type a name.
Everything above rests on choices that could have gone another way, and the honest version of this page is the one that shows how much the answer moves when they do. Four things are worth knowing before you trust any number here.
Sixty-five of these languages are in the database twice, described by different people working from different sources. That is an accident of collection, and it is also the only internal handle there is on how much of the apparent variety is language and how much is description.
Same-source pairs, where one wordlist has been entered twice (usually once in the source's spelling and once in phonetic transcription), are not independent and are counted separately. They agree more often, but not always, which puts a floor under the whole exercise: even copying one wordlist twice does not reliably preserve the system.
A language's system is the pattern of which relationships share a word, so everything turns on when two of these cells count as having the same word. The database gives strings. Some cells carry more than one string, because sources list variants, dialect forms and near-synonyms.
Three defensible rules, all computed:
A fourth choice turned out to matter more than it should, and it is the pettiest one on the page. When a language is in the database twice, one record has to be picked. The rule used here is: take the record with the fewest ambiguous cells, and break remaining ties by name and identifier, which is arbitrary but at least has nothing to do with the answer. The obvious improvement, preferring whichever record is written in the source's own spelling rather than in phonetic transcription, would make the page more readable, and it changes the count.
So that improvement was not adopted, because a rule that moves the answer is not a tiebreak, it is a choice about the answer wearing a tiebreak's clothes. Readability is handled at the other end instead: the words shown on the figures may come from a second record of the same language, but only ever one that yields the same system, so it cannot shift a count. That is why English reads brother and sister above rather than the 19th-century transcription sĭs'tər that the counting rule actually selected.
Spanish appears here with four sibling terms, because the source recorded hermano mayor, hermano menor, hermana mayor, hermana menor. Those are not four sibling words. They are two sibling words with an adjective attached, and a Spanish speaker would say so. Hindi is recorded as merá bará bhái, which carries a possessive as well: "my big brother".
The crude test for this is whether a recorded form contains a space, and 1,020 of the 10,748 forms here do. That test is used to keep such records out of the example lists, and it is wrong in both directions. It misses phrases written without a space. And it wrongly catches Mandarin, whose ge ge, di di, jie jie, mei mei really are four separate words that merely happen to romanise with a space in the middle.
The rule was left mechanical rather than corrected language by language, because a rule tuned per language is not a rule, it is a set of opinions with a script wrapped round it. So Mandarin is wrongly absent from the example lists, and this paragraph is the correction.
The eight cells and the number 4,140 come from Sara Nerlove and A. Kimball Romney, writing in American Anthropologist in 1967. We did not read them. The publisher returned a 403 and the archive copies are paywalled, so every statement here about that paper is second-hand, from sources that cite it, and those sources are quoted in full in the research directory. Two of them say Nerlove and Romney found 12 recurring types; a third says 23 systems attested in total, of which 15 occurred more than twice. We cannot adjudicate that from outside the article, so we report the disagreement and adopt none of the numbers.
What is confirmed, and quoted verbatim from the Kinbank paper itself, is the framing:
"These authors established a design space of 4,140 possible sibling terminologies, derived from an etic set of eight sibling categories derived from three rules of distinction (gender of sibling, gender of speaker, and relative age)."
4,140 is the eighth Bell number, the count of ways to partition an eight-element set, and the verifier recomputes it rather than taking anyone's word for it.
The sharper problem is a published count on this same database. In 2021 the team behind Kinbank wrote, in a footnote:
"Kinbank's larger sample contains a higher number of any-attested sibling types (116), of which four were more common than Nerlove and Romney's 8th and 9th commonest."
The likeliest explanation is that they were not counting on this release. Their footnote is from 2021; the database first reached Zenodo in 2022 and its descriptor paper in 2023, describing 1,229 languages and 210,903 terms where the release used here has 1,288 varieties and 218,654. A completeness rule we could not recover from a footnote would do it as well. We could not close the gap, so we are reporting it rather than adopting whichever number reads better.
One coincidence worth naming so that nobody mistakes it for a reconciliation: pooling every system that turns up under any of the three rules, across all 1,023 records rather than the 942 deduplicated languages, gives 117. That sits next to 116 and means nothing. It is a different quantity, computed a different way, and the closeness is luck.
Every combination of the three choices a recount has to make: which field of the database to read, whether a cell with no age-specific term may inherit the language's age-neutral one, and which of the three sameness rules to apply.
| field | age fallback | rule | languages | systems |
|---|
The licence is genuinely ambiguous and we did not resolve it. The Kinbank archive
contradicts itself: its LICENSE file is the full text of CC BY 4.0, while the
.zenodo.json beside it declares CC BY-NC-4.0, which is what Zenodo then displays. The
CLDF metadata carries no rights field at all and the paper's data-availability statement names no
licence for the data. We complied with the stricter of the two: this page is free, carries no
advertising, no trackers and nothing gated, and names the dataset and its authors here. We are
flagging the conflict rather than quietly picking the licence that suits us.
Citations, every verbatim quotation this page rests on, and the full list of what we did not
check are in research/sibling-partitions/SOURCES.md. The extraction is
research/sibling-partitions/extract.mjs and reruns from a fresh download of the
dataset.
The verifier is verify-nobody-speaks-the-other-four-thousand.mjs, and it runs
801 checks, all passing. It does not import the extraction: the partition logic is written
a second time, from the recorded word strings up, so that agreement between the two means
something. Run it with --mutate and four deliberately false claims are injected;
all four must go red, because a verifier that cannot fail is not checking anything. A separate
research/sibling-partitions/browser-check.mjs drives this page in a real browser
through 49 checks, including that the instrument can reach every one of the 4,140
arrangements by clicking, which is proved exhaustively rather than assumed.