Ground truth · the verification venue

The Dice Came Back Syllabic

In 2009, two journals published a statistic that was said to tell a writing system from a symbol system that is not writing. It placed the undeciphered Indus script among the languages, and it announced that Pictish symbols were "revealed as a written language." Here is the published rule, run on dice. Here it is run on chess, on a bacterial genome, on newspaper weather icons, and on fifteen languages. And here is the corpus that every Indus number in the argument rests on, which no one outside the original group has ever been able to see.

The Indus Valley civilisation left behind several thousand short inscriptions, most of them four or five signs long, on seals and tablets and pots. Nobody can read them. Nobody even agrees on the prior question: whether the signs encode a language at all, or whether they are something else, the way a coat of arms or a set of religious emblems is something else.

In May 2009 a group led by Rajesh Rao published a one-page piece in Science arguing the question could be settled statistically. Their instrument was conditional entropy: given one sign, how uncertain are you about the next one? Fully rigid systems score near zero, because the next sign is forced. Fully random systems score at the maximum, because nothing is forced. Natural languages, they argued, live in a characteristic middle band, and so does the Indus corpus. Ten months later, a paper in Proceedings of the Royal Society A went further and said Pictish symbols were a written language, by a related measure.

Richard Sproat, a computational linguist, replied in Computational Linguistics that the middle band was an artifact of what the papers had chosen to compare against. The authors replied to him. He replied to them. And there, in 2010, it stopped.

What nobody did was the obvious thing: take both published procedures and run them over a set of symbol systems whose answer is already known. If a measure can tell writing from not-writing, it has to sort English from chess. That is the whole of what is below.


I. The rule, and the dice

Lee, Jonathan and Ziman's Pictish paper builds two quantities and a decision tree. This is Sproat's transcription of them, which the three authors answered in print without disputing:

Ur = F2 / log2(Nd/Nu)  ·  Cr = Nd/Nu + a · Sd/Td, with a = 7

where F2 is the bigram entropy, Nd is the number of bigram types, Nu is the number of unigram types, Sd is the number of bigrams that occur once, and Td is the total number of bigram tokens. "If Cr ≥ 4.89, the system is linguistic. Subsequent refinements use values of Ur to classify the system as segmental (Ur < 1.09), syllabic (Ur < 1.37), or else logographic."

Richard Sproat, Computational Linguistics 36(3), 2010, pp. 590–591

Look at what the gate is made of before running it. The decision that a system is writing at all is Cr alone, and Cr contains no entropy: it is a count of distinct sign pairs divided by a count of distinct signs, plus the rate at which pairs occur exactly once. Entropy enters only at Ur, which merely sorts an already-admitted system into segmental, syllabic or logographic. The paper is titled Pictish symbols revealed as a written language through application of Shannon information theory.

Both of Cr's ingredients grow as you collect more text. Vocabulary accumulates; so does the stock of pairs you have seen once. Nothing in the rule mentions how much text you have.

Sproat answered the rule with noise. He built 75 short "texts" whose signs came from tossing seven six-sided dice, 638 symbols in all, and put them through the tree. It came back syllabic writing, with Cr = 12.64 and Ur = 1.18. He rolled it once. Roll it yourself, as often as you like.

Bench I · roll the dice through the published rule

638 signs, each the sum of seven six-sided dice, memoryless by construction

press roll


    

"Symbols derived by successive tosses of seven six-sided dice" does not say what a symbol is, and the readings give very different corpora, so the reading that produced his numbers had to be recovered. Six were tried, four thousand corpora each, built to his stated shape (75 texts, lengths 3 to 14, exactly 638 signs). Exactly one can reach Cr = 12.64: a sign is the sum of seven dice, which gives 36 possible values on a bell curve. That is also what his stated motive calls for, a source that is random but not equiprobable.

One half of his result does not reproduce, and it is worth saying exactly which. Cr = 12.64 is reachable and lands 62 times in 4,000 rolls. Ur = 1.18 is not: across all 24,000 corpora, over all six readings, the highest Ur seen is 1.149, and the two published values never occur together. So the corpora built here are called segmental writing about four times in five and syllabic writing the rest of the time, rather than syllabic every time. His own footnote records that the authors' published equation for bigram entropy "is apparently wrong", which would be enough to move Ur and not Cr; that is a plausible explanation and not a demonstrated one. None of it touches the part that matters. The decision that a system is writing at all is Cr, Cr reproduces, and every one of the 24,000 noise corpora passes it.

six readings of "seven six-sided dice", 4,000 corpora each

Rolling it once tells you noise can pass. Rolling it many times, across corpus sizes, tells you something worse.


II. Take the order away

Conditional entropy is supposed to measure order: how much the previous sign tells you about the next. There is an exact way to ask whether it is doing that. Every corpus has a twin with the same texts, the same lengths, and the same sign frequencies to the last count, but with the signs re-dealt at random. The twin has no order in it at all. Whatever number survives the shuffle was never about order.

This also solves the hard technical problem. The plug-in entropy estimator is biased downward, severely, when a corpus has few tokens per sign type, and the Indus corpus has about seventeen. But the bias falls on the twin too, and nearly equally, so the difference between them is close to unbiased. That difference is the order structure the estimator can actually see.

Each system below is cut to 7,000 signs, the size of the Indus corpus Rao et al. used. Pick one and take its order away.

Bench II · the same corpus, with and without its order


    

Take a language apart at word level and there is often nothing left to find. At 7,000 signs, French words carry · bits of detectable order and Hungarian words carry ·; the estimator cannot distinguish real prose from the same prose with its words thrown in a bag. Chess, over the same budget, carries · bits, several times more than any word-level language in the panel.

The reason is not that French has no word order. It is that at seventeen tokens per type, or three, an estimator built out of counts has almost nothing to count with. The whole panel, measured against its own shuffles:

Across the panel, a corpus's conditional entropy is predicted by the conditional entropy of the same corpus with its order destroyed to R2 = ·. Adding a variable that says whether the system is a language raises that to ·, a coefficient of · bits, which shuffling the labels reproduces about · of the time. The average system carries · bits of order against a spread across systems of · bits. Nearly all of what the statistic is reading is not order.


III. Rao's figure, with the controls it did not have

The Science paper's central figure plots conditional entropy against, in its own axis label, "number of tokens." Sproat records what the axis really is: the conditional entropy of the subset of each corpus made of the 20 most common signs, then the 40 most common, then the 60, and so on. That is a real procedure and it is run below on every system in the panel, including the ones the original figure did not have. The vertical axis is Rao et al.'s own normalisation, conditional entropy divided by the maximum a system of that inventory could reach, so that alphabets of different sizes share a scale.

The two systems that make the original figure look decisive are DNA and protein, and they are decisive because a genome really is close to memoryless at the level of single bases: its relative conditional entropy is ·, almost the ceiling. Nothing a human made on purpose sits up there. Chess does not. Newspaper weather icons do not. Barn stars, totem poles and Mesopotamian deity symbols do not. They sit in the band, because being made on purpose is what puts you in the band, and writing is not the only thing humans do on purpose.


IV. The better test, and why it does not help either

When Sproat objected that the method could not tell a random but non-equiprobable system from a language, Lee, Jonathan and Ziman answered with a genuinely better test, in their own words:

"For a given script set, the value of second-order entropy, E2, is calculated and compared with the corresponding value, E2(R), for a randomized script R consisting of a randomized permutation of the unigrams comprising the original script set. The probability P = Prob(E2(R) > E2) is estimated empirically using 1,000 randomized permutations. ... For approximately 80% of the script sets, the value of probability P is unity ... For both Pictish script sets, the estimated value of probability P is unity."

Lee, Jonathan and Ziman, Computational Linguistics 36(4), 2010, p. 792

That is the right test, and it is exactly the shuffle above. It is run here their way, 1,000 permutations per system, on every corpus in the panel.

P is a question about whether an effect is nonzero. With sixty thousand tokens, almost any real effect is nonzero, including the eighteen thousandths of a bit of neighbour dependence in a bacterial genome. What P never asks is how big the effect is, and the effect sizes here run from 0.018 bits to 1.7, in no order that tracks language. A criterion that a genome passes is not a criterion for writing.


V. Does anything here classify?

The panel holds · systems that are certainly writing and · that are certainly not. Every class was assigned before any number was computed and none of them is a close call. Scoring a measure by AUC asks: pick one writing system and one not-writing system at random, how often does the measure put them on the correct sides? A perfect measure scores 1.000. A coin scores 0.500.

And the single number this whole page turns on. Take Rao et al.'s own normalised statistic, at a matched sample size, and change only which non-linguistic systems it is asked to separate from:

Against DNA and protein, the controls the 2009 paper used, the separation is perfect. Against symbol systems that people made on purpose and that are not writing, it is worse than a coin. Both columns are shown because at a matched budget of 7,000 signs the three smallest human-made corpora (barn stars, totem poles, deity symbols) are too small to take part, so the matched column rests on three controls and the own-size column on six. The collapse is the same in both. Sproat's charge was that the band was an artifact of the comparison set. That is what an artifact of the comparison set looks like when you measure it.

The check

·

The estimator was anchored before anything was built on it: it reproduces F2 = 3.3006593072662582, the conditional entropy of one English letter given the previous one, already published by this archive at You Already Know the Rest, to zero bits on an identical 407,719-symbol stream (crosscheck-f2.mjs).

And the instrument was proved able to read zero. The whole shuffle procedure was run on corpora that have no order by construction, where the true answer is exactly 0.000 bits. Over · system-and-budget cells the measured value came out at · bits on average, largest magnitude ·.

Everything is in research/is-it-writing/: the corpora (committed, 4.6 MB gzipped), every build script with its conventions written down, the seeded analysis, and verify.mjs.

The wall: the corpus nobody can see

Every published Indus conditional entropy rests on a digitisation of Mahadevan's 1977 concordance: 1,548 lines of text, about 7,000 sign occurrences, 417 sign types. None of the numbers on this page is one of theirs, because that corpus cannot be obtained. On 2 August 2026 the routes were tried directly. The Science article page returns 403. Its Supporting Online Material describes the dataset and ships no data. The PNAS and PLOS ONE supplements carry no sign sequences. The arXiv source tarball for the companion paper holds 97 files, all figures and TeX. The database that holds the corpus, the Interactive Corpus of Indus Texts, returns HTTP 401, and its own landing page says: "Ask the administrator to get access to the database."

Sproat put this in print in 2014: "none of the Indus corpora that have been collected by various groups over the years have been made available to other researchers, which has in turn made it difficult to verify statistical results claimed for these corpora." Twelve years later it is still true. The Pictish side is the same shape: Lee, Jonathan and Ziman's source database was hosted at mathstat.strath.ac.uk, and that hostname no longer resolves.

So the finding about the Indus script is not that the claim is false. It is that the claim has never been checkable by anyone outside the group that made it, and has now outlived the machine it was computed on.