The Rule at the End of the Line
If the Voynich manuscript means nothing, something made it: a procedure, or a person making words up. Both kinds of maker exist and have been published. This page runs them through the same measures as the manuscript and fifty medieval books, and asks which of the manuscript's oddities any of them can reproduce.
Beinecke MS 408 is written in a script no one has read, on parchment radiocarbon dated to between 1404 and 1438. A companion layer measured its text beside fifty medieval books transcribed line by line and found it outside all fifty in two ways (its next letter is unusually easy to guess, and it writes the same word twice in a row more often than chance) and at the edge of them in a third: its words behave differently at the edges of a line, as among the fifty only verse, a printed list and one scribe's shorthand do. What that layer could not say is whether those measures tell a language from no language. Fifty meaningful books cannot answer that. You need texts that are known to mean nothing, and you need to put them through exactly the same code.
Here are five, and a sixth made by people:
- The self-citation generator of Torsten Timm and Andreas Schinner, published in Cryptologia in 2020 as
a concrete text-generator algorithm (the “self-citation” process), easily executable without additional tools even by a medieval scribe
. It copies words from lines already written and alters them. This is their own program (MIT licence), run unmodified: their default settings reproduce their published text byte for byte, and it was run here for twenty seeds at the manuscript's length. - The same program with one rule removed (section 3), again with its copying from the same place on the line turned up from 28% to 100%, and with a different rule removed (section 2).
- A word chain trained on the manuscript: each word is drawn from the words that followed the previous word somewhere in the manuscript. It knows which word follows which, and nothing about lines.
- The manuscript shuffled: its own words in random order, every line keeping its length.
- Latin in a verbose cipher: the twelve Latin books among the fifty, each letter replaced by a Voynich glyph group.
- Gibberish written by people, from Daniel Gaskell and Claire Bowern's 2022 experiment. Volunteers were told to
create a ‘document’ by filling three pages with fake, meaningless text in a ‘language’ that you make up as you go
, by hand, in pen. Their 38 public transcriptions are used here as released; 36 are long enough for the smallest window this page measures. A “writer” below means one sample: the volunteers wrote one each, and the authors contributed three.
Loading the manuscript’s transliteration, the generator’s texts and the comparison set… This page needs JavaScript to measure anything; the numbers in the prose were written from the same code and are checked by the verifier named at the end.
- The manuscript's predictable letters are a property of its words, not of their order. Shuffle its words into random order and h2 barely moves (2.10 against 2.03 bits); every machine that writes in its alphabet lands beside it. The number cannot count for or against meaning.
- The generator reproduces the manuscript's line endings, more strongly than the manuscript has them, because a ten-entry rule in its source code rewrites the ending of the word that closes a line, whenever that ending is in its table. Remove that rule and the line-end signal falls from 0.137 to 0.003. Turn copying-by-position up to 100% and it stays at 0.002. In this program the line ending is written in, not a side effect.
- No machine tried here reaches the manuscript's line starts: 0.149 in the manuscript, 0.043 to 0.059 in every run of the generator, with or without its rules, and about zero in the word chain and the shuffle. (The cipher keeps its Latin book's lines, and carries whatever line edges that book had.) In the people's gibberish, the median window carries about a tenth of the manuscript's excess at the line start.
- The same word twice in a row: the manuscript 1.46 times the shuffled rate; the generator 0.85 to 1.03, about chance, because a second rule in its code throws away half its repeats; with that rule taken out, 1.37 to 1.74, around the manuscript. The word chain overshoots, 2.03 to 2.70; the people's gibberish, all together, 36 against 62 expected, so it avoids the doubled word, though a few writers do it heavily.
- Twenty-seven predictions were committed before any of this was measured. Twenty-two held and five failed; two of the twenty-two held by the letter of the rule and for the wrong reason, and the page says which.
1. Which one did the scribe write?
Four texts, six consecutive lines each, in the EVA transliteration: one from the manuscript, one from the generator's published run, one from the word chain, one from the shuffle. The order is random. The chain and the shuffle were made in your browser a moment ago.
Whatever you score, the rest of this page does not rely on the eye.
2. Five measures, eight kinds of text
Each strip is one measure from the companion layer, computed by its own code (engine.mjs, imported here unchanged). Each dot is one run of one machine, or one book. The gold line is the manuscript, measured just now in your browser; the short white tick at its top is the same measure on a second, independent transliteration (Takeshi Takahashi's), to show how much the choice of transcriber moves it. The white rings are the runs your browser made or re-measured itself.
Tap or focus any dot to name it.
Letter predictability (h2). The manuscript's 2.03 bits was the companion layer's headline: lower than all fifty books. But the shuffle in your browser gives 2.10, and a shuffle has no order at all. The low figure comes from the words themselves, from the way each Voynich word is spelled out of a few glyphs in a nearly fixed order, and it survives any rearrangement of those words. Every machine that borrows the manuscript's spelling inherits it, meaningful or not. The one machine that writes real Latin through a cipher does not: my cipher gave the commonest Latin letters the commonest Voynich glyphs, which are single letters in EVA, so it was barely verbose and stayed near the books at 2.91 to 3.02. A cipher that spelled every letter with several glyphs would come lower; this one was not built to, and the prediction I made for it failed.
The same word twice. Each book writes identical neighbours less often than a shuffle of its own words would (at most 0.37 times); the manuscript writes them 1.46 times as often. The generator, which copies and alters nearby words, lands at about chance (its published run: 0.96), and its source code has a rule that throws away an exact repeat of the previous word half the time (commented // limit repeated words). Yet it matches the manuscript almost exactly on neighbours one letter apart (1.44 against 1.47). It makes near-copies at the manuscript's rate and exact copies at a shuffle's. So I took that rule out too (a second one-line change: keep every repeat instead of half), leaving the line-end rule in place, and ran twenty more seeds. The doubled words came back: 1.37 to 1.74 times the shuffled rate, above the manuscript's in 17 of the 20. Copying nearby words makes the manuscript's doubled words by itself; the published generator had been built to suppress them. The word chain, which reproduces every pair the manuscript contains, overshoots to more than twice chance. That is a property of the yardstick: the shuffle it is compared with mixes words only within runs of 30 lines. The manuscript's vocabulary changes from page to page, so a local shuffle already produces many of its repeats; the chain spreads every word evenly through the book, so a local shuffle cannot. Against a shuffle of the whole text instead, the manuscript scores 2.87 and the chain 1.97 to 2.71 over twenty seeds, every one below it.
The edges of the line. These two strips separate the machines most sharply. Anything that does not know where its lines end scores about zero: the chain and the shuffle. The cipher keeps each Latin book's own lines, so it carries that book's line edges, low for most of them. The generator scores higher than the manuscript at the line end, and a third of the manuscript at the line start. Section 3 is why.
3. The rule at the end of the line
In 1976 the cryptanalyst Prescott Currier told a seminar that the line is a functional entity
in the manuscript, because the frequency counts of the beginnings and endings of lines are markedly different from the counts of the same char-acters internally
. The companion layer counted it: words ending in m stand last on their line 70% of the time, words ending in g 79%, where the average word is last 12% of the time.
In 2014 Torsten Timm offered an explanation in which the line ending is an accident of the method:
Both observations can be explained as an unintended side effect of the text generation method. The source for the first word in each line could only be found within the previous lines. Since the first and the last word in each line are easy to spot, the most obvious way is to pick them as a source for the generation of a group at the beginning or at the end of a line.T. Timm, “How the Voynich Manuscript was created”, arXiv:1407.6639 (v3, 2015), p. 19
The generator published five years later does copy from the same place on earlier lines (with probability 28% by default). But it also contains this table, in Glyph.java. The main loop consults it whenever the next word will not fit in the space left on the line, which is to say for the word that ends the line, and rewrites that word's ending by it, under the comment // change last token 'ol' --> 'om':
So I switched that one function off (a one-line change, recompiled against the authors' own compiled program; an unmodified recompile reproduces their published text byte for byte) and ran it again, twenty seeds. Then again with copying-by-position at its maximum. Here are the line ends, measured in your browser:
Line-end information: the manuscript 0.089, the generator as published 0.142, the generator with the rule removed 0.002. “None”: no word in that text ends in the letter; “rare”: fewer than 20 do.
With the rule gone, m and g vanish from the generator's text altogether, and with them its whole line-end signal: 0.127 to 0.150 over twenty seeds with the rule, 0.002 to 0.005 without, 0.000 to 0.005 without it and with copying from the same place on an earlier line turned up to its maximum. In this program, copying by position does not carry the line ending, because nothing puts one there for it to copy. The rule does.
That does not refute the hoax. A fifteenth-century forger could have kept a rule like this as easily as a program can: write the last word of the line differently. It changes what the generator's success at the line end is evidence of. It is not a property that self-citation produces; it is a property the authors observed in the manuscript and wrote into their machine, which is a fair thing to do and is also, for the purpose of telling a hoax from a language, circular. The same is true of language: rhyme gives verse a line ending, and one Castilian scribe in the fifty kept a raised s almost entirely for line ends. Something in the maker knew where the line was. Every maker tried here that did not know, scores zero.
The line start is the harder case. The generator has rules there too (gallows letters for the first word of a paragraph, an o, y, d or s added in front of a line's first word three times in ten, a different copying probability for that word) and with them it reaches 0.057. The manuscript's is 0.149.
4. People making it up
Gaskell and Bowern's volunteers were mostly students on a Yale course about the manuscript (the 2018 class wrote before they had studied its statistics, the 2019 class later in the semester), plus three members of the public who did not know the study's connection to it, plus the authors. Their conclusion was that gibberish varies widely and, depending on the sample, is able to replicate either natural language or Voynichese across nearly all of the metrics which we tested
, and that the results refute the idea that the low-level linguistic structure of the VMS text is too non-random to be meaningless
. They also wrote that our writing samples are too short to test whether the higher-level structure of VMS pages and quires could also be produced by gibberish
.
Their samples run from 79 to 480 words, too short for the measures above, which take h2 over stretches of 20,000 letters and the rest over the whole text. So here every text is cut the same way into windows of whole lines of at least 150 words, and each window is measured on its own. At that size the line-edge estimate is biased upward even for text that ignores its lines, so each window is compared with itself shuffled: the figure plotted is how far the window's line-edge information stands above its own shuffled floor. (That correction was added after the first run, when the line-blind word chain scored 0.05 to 0.09 here; it was not in the pre-registration.)
The people write in the Latin alphabet and the manuscript is read through EVA, so their h2 values are not on the same footing; the books share the people's alphabet. The generator row is one run (seed 1), which gives 248 windows; the books row plots a thinned sample of 5,200 windows, and its median uses all of them.
On h2 the people land among the books, not below them (median 2.97 bits against 2.84 for book windows), and I had predicted the opposite. At the line edges, the median window of gibberish stands 0.013 above its floor at the start and 0.011 at the end; book windows 0.011 and 0.021; the manuscript's windows 0.140 and 0.059. Two of the 36 samples reach the manuscript's median at the line start over their one to three windows, which is a small and noisy sample; one of the two is, by the dataset's own metadata, an author's.
And the same word twice. Gaskell and Bowern report that gibberish repeats words more than meaningful text does, and it does: the books' windows write identical neighbours at 0.06 times their shuffled rate, the people at 0.58. But against its own shuffle, the people's gibberish still avoids the doubled word, where the manuscript seeks it:
| Writer | Words | Neighbour pairs | Same word twice | Expected if shuffled |
|---|
Nine of the 36 samples double words more often than their shuffle would; 23 never do it at all. The strongest, DC_20, wrote nine doubled words where 1.6 were expected. The manuscript at this scale: 287 against 219. What people do when inventing text is not one thing, which is Gaskell and Bowern's first finding, and on this measure the manuscript is nearer the unusual samples than the typical one.
5. What was predicted, and what failed
The predictions were committed to the repository before any machine was run (PREREGISTRATION.md). The rule: a kind of text lands with the manuscript on a measure if its median is nearer the manuscript's value than the nearest of the fifty books; otherwise with the books.
| Text | Measure | Predicted | Result | |
|---|---|---|---|---|
| generator | h2 | manuscript | 2.010, manuscript | held |
| generator | same word twice | manuscript | 0.982, manuscript | held by the letter |
| generator | one letter apart | manuscript | 1.452, manuscript | held |
| generator | line end | books | 0.137, books | held, wrong reason |
| generator | line start | manuscript | 0.053, books | failed |
| word chain | h2, same word twice | manuscript | 2.096, 2.285, manuscript | held |
| word chain | one letter apart | manuscript | 1.945, books | failed |
| word chain | line end, line start | below 0.02 | 0.001, 0.001 | held |
| cipher | h2 | 2.0 to 2.8 bits | 2.970 | failed |
| cipher | the other four | books | books, all four | held |
| shuffle | all five | h2 manuscript; ratios 0.9 to 1.1; line edges below 0.01 | 2.100; 1.047, 0.989; 0.001, 0.001 | held |
| people | same word twice, pooled | above 1 | 0.577 | failed |
| people | window h2 | below the books' median | 2.966 against 2.836 | failed |
| second transliteration | all five | within 10% of ZL | largest difference 4.7% | held |
Twenty-seven predictions, counting each measure separately: 22 held, 5 failed. Two of the holds deserve no credit. The generator's doubled words (0.982) sit at chance, far from the manuscript, and count as "manuscript" only because the nearest book is further still; the rule was too blunt for a value that lands between. And I predicted the generator would miss the line end because, reading its main file, I found a special case for the first word of a line and none for the last. The special case for the last is on line 321 of the same file. The prediction held only because the generator overshoots the manuscript so far that one of the books is nearer: the Castilian prose manuscript whose scribe kept a raised s for line ends.
Three things were added after the first results and are not pre-registered: the generator with its line-end rule removed (and with copying by position at 100%), the generator with its repeat rule removed, and the shuffled floor for line edges at window size. A table-and-grille generator after Gordon Rugg (2004) was pre-registered only if it could be built from his paper's own description; the paper could not be obtained, so it was not built.
What this cannot tell you
Whether the manuscript is a hoax. Every text here is one maker. A different generator, a cipher that spells each letter with several glyphs, a writer with a habit of doubling words: any of them might land nearer. What the measures can do is say which of the manuscript's oddities come free with a way of making text and which have to be put there on purpose. Its letter statistics come free with its spelling. Its line endings have to be put there, by a rule or by a convention, and the one published machine that has them had them put there. Its doubled words come free with self-citation, once the rule that suppresses them is taken out. Its line starts, nothing here produced. That is a list of targets for the next generator, and a list of things a proposed language would have to explain. Montemurro and Zanette (2013) argued from the way the manuscript's words cluster by section that it has a complex organization in the distribution of words that is compatible with those found in real language sequences
; Timm replied that the context dependency only points to a self-referencing system
. This page measures neither claim.
The people were mostly students on a course about the manuscript, writing a few hundred words each, in a Latin alphabet. The manuscript's running text is about 35,000 words. There is no data here on what a person inventing text does at that length.
The check
- Measured as this page loads, by
machines.mjs, which imports the companion layer'sengine.mjsunmodified: the manuscript (ZL 3b, CC0, from voynich.nu), Timm and Schinner's published run (sc-published-seed19.txt, byte-identical to their repository's file), the generator with its line-end rule removed (seed 19), and one run each of the word chain and the shuffle. - Every other dot is in
data/machines.json, written by the scripts inresearch/voynich-machines/:gen-sc.shfetches the authors' jar, checks that an unmodified recompile ofGlyph.javareproduces their published text byte for byte, and runs the three generator families;run-long.mjsandrun-short.mjsmeasure everything;score.mjsscores the predictions. The fifty books' long-scale figures are the companion layer's own file; their windows were re-measured from CATMuS Medieval (all 194,808 rows, text columns only). - The people's gibberish is
data/gibberish.json, the 38 transcriptions in Gaskell and Bowern's public repository, unaltered, under their modified MIT licence, which asks that the paper be cited. It is, below. node verify-is-the-voynich-manuscript-a-hoax.mjsre-measures the served texts, recomputes every figure the prose states from the data, re-scores all 27 predictions, and checks that the measures respond as they should: a generator text with the rule re-applied by hand regains its line end, and a text with every line's first word moved to its end loses its line start.