The Rule at the End of the Line

If the Voynich manuscript means nothing, something made it: a procedure, or a person making words up. Both kinds of maker exist and have been published. This page runs them through the same measures as the manuscript and fifty medieval books, and asks which of the manuscript's oddities any of them can reproduce.

Beinecke MS 408 is written in a script no one has read, on parchment radiocarbon dated to between 1404 and 1438. A companion layer measured its text beside fifty medieval books transcribed line by line and found it outside all fifty in two ways (its next letter is unusually easy to guess, and it writes the same word twice in a row more often than chance) and at the edge of them in a third: its words behave differently at the edges of a line, as among the fifty only verse, a printed list and one scribe's shorthand do. What that layer could not say is whether those measures tell a language from no language. Fifty meaningful books cannot answer that. You need texts that are known to mean nothing, and you need to put them through exactly the same code.

Here are five, and a sixth made by people:

Loading the manuscript’s transliteration, the generator’s texts and the comparison set… This page needs JavaScript to measure anything; the numbers in the prose were written from the same code and are checked by the verifier named at the end.

What the measuring finds.

1. Which one did the scribe write?

Four texts, six consecutive lines each, in the EVA transliteration: one from the manuscript, one from the generator's published run, one from the word chain, one from the shuffle. The order is random. The chain and the shuffle were made in your browser a moment ago.

Whatever you score, the rest of this page does not rely on the eye.

2. Five measures, eight kinds of text

Each strip is one measure from the companion layer, computed by its own code (engine.mjs, imported here unchanged). Each dot is one run of one machine, or one book. The gold line is the manuscript, measured just now in your browser; the short white tick at its top is the same measure on a second, independent transliteration (Takeshi Takahashi's), to show how much the choice of transcriber moves it. The white rings are the runs your browser made or re-measured itself.

Letter predictability (h2). The manuscript's 2.03 bits was the companion layer's headline: lower than all fifty books. But the shuffle in your browser gives 2.10, and a shuffle has no order at all. The low figure comes from the words themselves, from the way each Voynich word is spelled out of a few glyphs in a nearly fixed order, and it survives any rearrangement of those words. Every machine that borrows the manuscript's spelling inherits it, meaningful or not. The one machine that writes real Latin through a cipher does not: my cipher gave the commonest Latin letters the commonest Voynich glyphs, which are single letters in EVA, so it was barely verbose and stayed near the books at 2.91 to 3.02. A cipher that spelled every letter with several glyphs would come lower; this one was not built to, and the prediction I made for it failed.

The same word twice. Each book writes identical neighbours less often than a shuffle of its own words would (at most 0.37 times); the manuscript writes them 1.46 times as often. The generator, which copies and alters nearby words, lands at about chance (its published run: 0.96), and its source code has a rule that throws away an exact repeat of the previous word half the time (commented // limit repeated words). Yet it matches the manuscript almost exactly on neighbours one letter apart (1.44 against 1.47). It makes near-copies at the manuscript's rate and exact copies at a shuffle's. So I took that rule out too (a second one-line change: keep every repeat instead of half), leaving the line-end rule in place, and ran twenty more seeds. The doubled words came back: 1.37 to 1.74 times the shuffled rate, above the manuscript's in 17 of the 20. Copying nearby words makes the manuscript's doubled words by itself; the published generator had been built to suppress them. The word chain, which reproduces every pair the manuscript contains, overshoots to more than twice chance. That is a property of the yardstick: the shuffle it is compared with mixes words only within runs of 30 lines. The manuscript's vocabulary changes from page to page, so a local shuffle already produces many of its repeats; the chain spreads every word evenly through the book, so a local shuffle cannot. Against a shuffle of the whole text instead, the manuscript scores 2.87 and the chain 1.97 to 2.71 over twenty seeds, every one below it.

The edges of the line. These two strips separate the machines most sharply. Anything that does not know where its lines end scores about zero: the chain and the shuffle. The cipher keeps each Latin book's own lines, so it carries that book's line edges, low for most of them. The generator scores higher than the manuscript at the line end, and a third of the manuscript at the line start. Section 3 is why.

3. The rule at the end of the line

In 1976 the cryptanalyst Prescott Currier told a seminar that the line is a functional entity in the manuscript, because the frequency counts of the beginnings and endings of lines are markedly different from the counts of the same char-acters internally. The companion layer counted it: words ending in m stand last on their line 70% of the time, words ending in g 79%, where the average word is last 12% of the time.

In 2014 Torsten Timm offered an explanation in which the line ending is an accident of the method:

Both observations can be explained as an unintended side effect of the text generation method. The source for the first word in each line could only be found within the previous lines. Since the first and the last word in each line are easy to spot, the most obvious way is to pick them as a source for the generation of a group at the beginning or at the end of a line.T. Timm, “How the Voynich Manuscript was created”, arXiv:1407.6639 (v3, 2015), p. 19

The generator published five years later does copy from the same place on earlier lines (with probability 28% by default). But it also contains this table, in Glyph.java. The main loop consults it whenever the next word will not fit in the space left on the line, which is to say for the word that ends the line, and rewrites that word's ending by it, under the comment // change last token 'ol' --> 'om':

So I switched that one function off (a one-line change, recompiled against the authors' own compiled program; an unmodified recompile reproduces their published text byte for byte) and ran it again, twenty seeds. Then again with copying-by-position at its maximum. Here are the line ends, measured in your browser:

With the rule gone, m and g vanish from the generator's text altogether, and with them its whole line-end signal: 0.127 to 0.150 over twenty seeds with the rule, 0.002 to 0.005 without, 0.000 to 0.005 without it and with copying from the same place on an earlier line turned up to its maximum. In this program, copying by position does not carry the line ending, because nothing puts one there for it to copy. The rule does.

That does not refute the hoax. A fifteenth-century forger could have kept a rule like this as easily as a program can: write the last word of the line differently. It changes what the generator's success at the line end is evidence of. It is not a property that self-citation produces; it is a property the authors observed in the manuscript and wrote into their machine, which is a fair thing to do and is also, for the purpose of telling a hoax from a language, circular. The same is true of language: rhyme gives verse a line ending, and one Castilian scribe in the fifty kept a raised s almost entirely for line ends. Something in the maker knew where the line was. Every maker tried here that did not know, scores zero.

The line start is the harder case. The generator has rules there too (gallows letters for the first word of a paragraph, an o, y, d or s added in front of a line's first word three times in ten, a different copying probability for that word) and with them it reaches 0.057. The manuscript's is 0.149.

4. People making it up

Gaskell and Bowern's volunteers were mostly students on a Yale course about the manuscript (the 2018 class wrote before they had studied its statistics, the 2019 class later in the semester), plus three members of the public who did not know the study's connection to it, plus the authors. Their conclusion was that gibberish varies widely and, depending on the sample, is able to replicate either natural language or Voynichese across nearly all of the metrics which we tested, and that the results refute the idea that the low-level linguistic structure of the VMS text is too non-random to be meaningless. They also wrote that our writing samples are too short to test whether the higher-level structure of VMS pages and quires could also be produced by gibberish.

Their samples run from 79 to 480 words, too short for the measures above, which take h2 over stretches of 20,000 letters and the rest over the whole text. So here every text is cut the same way into windows of whole lines of at least 150 words, and each window is measured on its own. At that size the line-edge estimate is biased upward even for text that ignores its lines, so each window is compared with itself shuffled: the figure plotted is how far the window's line-edge information stands above its own shuffled floor. (That correction was added after the first run, when the line-blind word chain scored 0.05 to 0.09 here; it was not in the pre-registration.)

On h2 the people land among the books, not below them (median 2.97 bits against 2.84 for book windows), and I had predicted the opposite. At the line edges, the median window of gibberish stands 0.013 above its floor at the start and 0.011 at the end; book windows 0.011 and 0.021; the manuscript's windows 0.140 and 0.059. Two of the 36 samples reach the manuscript's median at the line start over their one to three windows, which is a small and noisy sample; one of the two is, by the dataset's own metadata, an author's.

And the same word twice. Gaskell and Bowern report that gibberish repeats words more than meaningful text does, and it does: the books' windows write identical neighbours at 0.06 times their shuffled rate, the people at 0.58. But against its own shuffle, the people's gibberish still avoids the doubled word, where the manuscript seeks it:

Nine of the 36 samples double words more often than their shuffle would; 23 never do it at all. The strongest, DC_20, wrote nine doubled words where 1.6 were expected. The manuscript at this scale: 287 against 219. What people do when inventing text is not one thing, which is Gaskell and Bowern's first finding, and on this measure the manuscript is nearer the unusual samples than the typical one.

5. What was predicted, and what failed

The predictions were committed to the repository before any machine was run (PREREGISTRATION.md). The rule: a kind of text lands with the manuscript on a measure if its median is nearer the manuscript's value than the nearest of the fifty books; otherwise with the books.

TextMeasurePredictedResult
generatorh2manuscript2.010, manuscriptheld
generatorsame word twicemanuscript0.982, manuscriptheld by the letter
generatorone letter apartmanuscript1.452, manuscriptheld
generatorline endbooks0.137, booksheld, wrong reason
generatorline startmanuscript0.053, booksfailed
word chainh2, same word twicemanuscript2.096, 2.285, manuscriptheld
word chainone letter apartmanuscript1.945, booksfailed
word chainline end, line startbelow 0.020.001, 0.001held
cipherh22.0 to 2.8 bits2.970failed
cipherthe other fourbooksbooks, all fourheld
shuffleall fiveh2 manuscript; ratios 0.9 to 1.1; line edges below 0.012.100; 1.047, 0.989; 0.001, 0.001held
peoplesame word twice, pooledabove 10.577failed
peoplewindow h2below the books' median2.966 against 2.836failed
second transliterationall fivewithin 10% of ZLlargest difference 4.7%held

Twenty-seven predictions, counting each measure separately: 22 held, 5 failed. Two of the holds deserve no credit. The generator's doubled words (0.982) sit at chance, far from the manuscript, and count as "manuscript" only because the nearest book is further still; the rule was too blunt for a value that lands between. And I predicted the generator would miss the line end because, reading its main file, I found a special case for the first word of a line and none for the last. The special case for the last is on line 321 of the same file. The prediction held only because the generator overshoots the manuscript so far that one of the books is nearer: the Castilian prose manuscript whose scribe kept a raised s for line ends.

Three things were added after the first results and are not pre-registered: the generator with its line-end rule removed (and with copying by position at 100%), the generator with its repeat rule removed, and the shuffled floor for line edges at window size. A table-and-grille generator after Gordon Rugg (2004) was pre-registered only if it could be built from his paper's own description; the paper could not be obtained, so it was not built.

What this cannot tell you

Whether the manuscript is a hoax. Every text here is one maker. A different generator, a cipher that spells each letter with several glyphs, a writer with a habit of doubling words: any of them might land nearer. What the measures can do is say which of the manuscript's oddities come free with a way of making text and which have to be put there on purpose. Its letter statistics come free with its spelling. Its line endings have to be put there, by a rule or by a convention, and the one published machine that has them had them put there. Its doubled words come free with self-citation, once the rule that suppresses them is taken out. Its line starts, nothing here produced. That is a list of targets for the next generator, and a list of things a proposed language would have to explain. Montemurro and Zanette (2013) argued from the way the manuscript's words cluster by section that it has a complex organization in the distribution of words that is compatible with those found in real language sequences; Timm replied that the context dependency only points to a self-referencing system. This page measures neither claim.

The people were mostly students on a course about the manuscript, writing a few hundred words each, in a Latin alphabet. The manuscript's running text is about 35,000 words. There is no data here on what a person inventing text does at that length.

The check