Fifty Books and One Unread One

Nobody can read the Voynich manuscript. Anyone can measure it. Here your browser measures the whole transliteration and sets it beside fifty medieval books, most of them copied by hand, all transcribed line by line, to ask a narrower question than “what does it say”: what kind of writing does it behave like?

Beinecke MS 408 is a small illustrated book of plants, stars, bathing figures and jars, written throughout in a script no one has read. Its parchment was radiocarbon dated in 2009 to between 1404 and 1438 (at 95% probability), which dates the parchment, not the ink. For a century people have tried to decipher it and failed. This page does not try. It asks what can be counted in the text, and what each count can and cannot tell apart.

The usual comparison is with modern printed books, and printed books are the wrong control. A printer breaks a line wherever the type runs out, a modern editor expands every abbreviation, and nobody decides at the end of a line to write a letter differently. A medieval scribe did all of those things. So the comparison set here is CATMuS Medieval: books from the ninth to the sixteenth century, transcribed line by line as they were written, abbreviations and all. Every one with at least 5,000 words of main text is used: 50 books, about 800,000 words, in Latin, French, Castilian, Middle Dutch and Italian, prose and verse. Forty-two are manuscripts; eight are early printed books, which CATMuS includes because their type imitates a scribe's hand and keeps his abbreviations. A switch below hides the printed ones.

Loading the transliteration (412 KB) and the comparison set… This page needs JavaScript to measure anything; the numbers in the text below were written from the same code and are checked by the verifier named at the end.

What the measuring finds.

1. How easy is the next letter to guess?

Take a long run of text, look at one letter, and ask how uncertain the next one is. That uncertainty, averaged over the text, is the second-order conditional entropy, h2, in bits. English, Latin, French: a little over three bits. The Voynich has been known since William Bennett's computer study of 1976 to sit far below that; he called it fantastically low, and gave 2.22 against 3.01 to 3.37 for ordinary languages. In 2020 Claire Bowern and Luke Lindemann compared it with 311 languages from Wikipedia and found none as low: the nearest was Hawaiian at 2.77.

There is a trap in this number. The Voynich is read through a transliteration, a Latin-letter code for its glyphs, and the standard one, EVA, writes some single-looking shapes as two or three letters: the “bench” glyph that looks like a joined pair of cs is ch, and a tall “gallows” glyph on a bench is cth. Splitting one glyph into two predictable letters makes the text look more predictable than it is. So the instrument below lets you choose how to cut it: EVA as written; with the six benched glyphs as single symbols; and also with the runs of i and their final stroke (in, iin, ir…) as single symbols, the way Currier's older alphabet did.

However generously the script is cut, the Voynich stays below every one of the fifty. Merging the benches and i-runs raises it from to bits and closes about a third of the gap to the most predictable book, a thirteenth-century French verse manuscript at 3.06. Abbreviations do not rescue the comparison either: a scribe who writes dñs for dominus adds symbols to the alphabet rather than taking them away.

What the number cannot tell you is why. Low h2 is what a language with few, rigidly ordered sounds produces (that is why Hawaiian comes nearest); it is also what a verbose cipher produces, one that spells each plain letter as a fixed group of symbols; and it is what a text built by copying and altering nearby words produces. A conditional entropy never settled whether a script was writing, and it does not settle it here.

2. Where a word sits on its line

Prescott Currier, a cryptanalyst who had served in the US Navy from 1929 to 1962, told a seminar in 1976 that he had noticed something he could not explain:

The first point is that the line is a functional entity in the manuscript on all those pages where the text is presented linearly. […] The frequency counts of the beginnings and endings of lines are markedly different from the counts of the same characters internally.Prescott Currier, “Some Important New Statistical Findings”, in New Research on the Voynich Manuscript (1976), transcribed at voynich.nu

Here it is, counted. On average a word is last on its line of the time. But look at the words ending in m or g, and at the words beginning with the gallows p or t:

Some lines from the manuscript, with the line-final words in m marked. None of these is the first or last line of its paragraph:

Currier said one symbol occurs at the end of the last words of lines 85% of the time, without naming it. In this transliteration the closest are words ending in g () and in m (). None of this is a paragraph effect: dropping the first and last line of every paragraph leaves it just as strong (line end , line start on the scale below).

To compare it with the fifty books, put it on one scale: how much does the last letter of a word tell you about whether the word ends a line, as a fraction of the uncertainty there is to remove? Zero means the line is invisible to the word. In prose printed by a printer it is close to zero. For the Voynich it is at the end of the line and at the start (the two middle strips above).

And here the scribes are more interesting than expected, because real scribes also wrote differently at the edge of a line. Three of the fifty reach or pass the Voynich at the line's end, all of them handwritten, and each for a reason you can see:

At the start of the line only one of the fifty passes the Voynich, and it is not running prose at all: a Latin book printed in the fifteenth century (Paris, Mazarine, Inc. 59) where much of the transcribed text is chapter headings, one to a line; 704 of its 1,880 lines begin De (“On…”). A list is a text whose line is a unit.

So the line effect does not make the Voynich unlike all writing. It makes it unlike prose. By this measure a Voynich line behaves like a line of verse, a line-end abbreviation habit, or an entry in a list.

3. The same word twice

Languages avoid saying the same word twice in a row. It happens (“that that”, “very very”), but less often than if the words of a page were shuffled into random order. Measure the rate of identical neighbours inside a line, divide by the rate after shuffling the words among a run of 30 lines, and every book in the comparison set comes out below one. The Voynich comes out at : of adjacent pairs are the same word, as in qokeedy qokeedy or chol chol (the strip above). Shuffling within each page instead of each run of 30 lines, which is harsher because a page's vocabulary is narrower, still gives .

Currier made a sharper claim, and it can be checked on a better transcription than he had:

I have three computer runs of the herbal material and of the biological material. In all of that, which is almost 25,000 words, there is not one single case of a repeat going over the end of a line to the beginning of the next; not one.Currier 1976, the same talk

Not none: six, all in the herbal and biological pages. But the inside of a line would have produced about twenty, and the rest of the manuscript would have produced about eleven and has none. Currier's “not one” was almost right, and his conclusion stands: repetition is something that happens within a line.

The last strip measures a looser kind of echo: neighbours one letter apart (qokeedy qokedy). Torsten Timm and Andreas Schinner proposed in 2020 that the text was made by exactly this, copying nearby words with small changes, and the Voynich is well above chance on it (). But this is where one of the predictions failed. Three Latin grammatical treatises of the ninth and tenth centuries reach 2 or more, because a grammar is full of paradigms set side by side (turba turbo, scintillo uacillo, uolo as), and short tokens like in, i and io are one letter apart by construction. Near-neighbours are not a fingerprint of generated text. A grammar book does it more.

4. Two “languages”, found blind

Currier's other discovery was that the manuscript is written in two statistically distinct varieties, which he called A and B while insisting the word in no way implies the existence of any underlying language. In 2020 the palaeographer Lisa Fagin Davis identified five scribes by their handwriting, and reported that scribe 1 writes A and the other four write B. The transliteration used here records both labels for each page.

The test: describe every page with at least 60 words by the frequencies of its 60 commonest letter pairs, and split the pages into two groups by k-means, never showing it a label. Then, and only then, compare the groups with Currier's.

The blind split agrees with Currier on pages. It is not just separating the plant pages from the rest: the herbal pages exist in both varieties, and of them land on Currier's side. The five disagreements are three pharmaceutical pages (f88v, f99v, f102v2) and the two sides of folio 58, which Currier's label calls A but which Davis gives to scribe 3, whose other labelled pages are all B.

The scribes themselves are harder to see. Guessing each page's scribe from the average profile of the other pages gets right, where always guessing scribe 1 would get . But hold the section fixed, the herbal pages in B, and it gets , where always guessing scribe 2 gets . The letters see Currier's two varieties sharply and Davis's scribes only faintly. That is consistent with Davis's reading, that A and B are habits of whole groups of writers, and it means these counts cannot confirm her five hands, which rest on the shapes of the letters, not their frequencies.

5. Measure your own text

The same five measurements, on anything you paste, with your line breaks kept. Try a poem against a paragraph of prose, or load one of the scribes. The samples are 400 lines each from CATMuS (CC BY 4.0), too short for the entropy chunks the manuscripts get, so their h2 is measured in one piece and reads a little low.

6. What was predicted, and what failed

The comparison set arrived slowly, and while it was still downloading the measures were run on the first 21 books that had enough text. Before measuring the rest, five predictions were written down and committed (PREREGISTRATION.md, which also records a correction to its own row count). They were scored on the 29 books not seen (the predictions said “manuscript” for all of them, printed ones included):

Prediction for the unseen 29Result
None has h2 below 2.60 bitsheld lowest 3.06
No prose manuscript reaches the Voynich's line-end value (0.089)held the Castilian scribe's 0.132 was among the first 21, already seen
No manuscript reaches its line-start value (0.149)failed the printed table of contents, Mazarine Inc. 59, at 0.207
No manuscript repeats a word at or above chanceheld highest 0.37
No manuscript has one-letter neighbours at or above chancefailed six do, three grammatical treatises at 2.00 to 2.16

The failures taught more than the successes: both of them are texts whose lines or neighbours are not running sentences. That is the thread through the whole page.

What this cannot tell you

None of these numbers reads a word, and none of them settles whether the text means something. They narrow what kind of writing it is like. On every measure here, the medieval books that come nearest to the Voynich are the ones whose lines are not prose: verse, lists, paradigms, a scribe's line-end shorthand. A natural language written in some such form, a cipher that works line by line, and a text generated by copying and varying nearby words would all leave marks of this general kind; the measures on this page do not choose between them. Montemurro and Zanette (2013) found that the way its words cluster across sections looks like a real text's; Timm and Schinner (2020) argued the opposite from the similarity of neighbouring words. What this page adds is the right control. Compared with medieval books as they were written rather than modern printed editions, the Voynich's odd lines are less alien than they looked, and its predictability and its repetitions are more.

Caveats, all of them. The Voynich is measured through one transliteration, Zandbergen and Landini's ZL (version 3b), paragraph text only, words with an unreadable or rare glyph left out; another transliteration would move every number a little. CATMuS writes letters in their modern forms (a long ſ is transcribed s), so any habit of a scribe that lived only in a letter's shape is invisible to it, while EVA records shapes; the scribes' line effects are if anything understated. CATMuS lines are not in page order, which does not affect anything measured here but means no measure here crosses from one line to the next except in the Voynich. Each book is whatever part of it CATMuS transcribed. Eight of the fifty are early printed books, not manuscripts; the switch in section 1 hides them. The page map depends on the choice of 60 letter pairs and 50 random starts; the verifier re-runs it.

The check