Cicada 3301 · the Liber Primus · 12,956 runes

One Bit of Structure

In 2014 Cicada 3301 published a book in runes and then went quiet. Nine sections of it have opened. Pages 0 to 55 have resisted a decade of public attack. Measured end to end, those pages depart from random noise in exactly one way, and that one way rules out almost everything anyone has tried.

The wall

This is the real thing: the first stretch of the unsolved pages, exactly as they sit in the book. Twenty-nine runes, each with a prime attached, arranged in words with real sentence punctuation. Nobody knows what it says.

Press the middle button, then the third. In the book, a rune is followed by itself 86 times in 12,955 chances. Shuffle the very same runes and it happens about 447 times. That gap is the entire subject of this page.

Everything else is noise

Before the gap means anything, you have to know it is alone. It is. Here is the bigram table, all 841 cells: how often each rune is followed by each other rune. English has visible structure. A Vigenere of English does not, but its diagonal is as bright as everything else. The book's diagonal is a hole.

Rows are the first rune, columns the second, brightness is frequency. The diagonal is "a rune followed by itself". Only one of these four has a dark one.

The rest of the measurements are as empty as the off-diagonal cells. The 29 rune frequencies are uniform (chi-square 26.4 on 28 degrees of freedom, which is exactly what uniform looks like). The repeat rate at every lag from 2 to 87 sits within three standard deviations of chance. No key length from 2 to 150 does better than a text with no message in it. The longest repeated run is six runes, which is what 12,956 random runes give you. Word sums are prime at exactly the rate chance predicts.

One statistic in the whole book departs from noise, and it is the one that says a rune will not follow itself.

Why that kills the Vigenere

Here is the argument, and you can check every number in it. Suppose the book is enciphered the way its own solved pages are: add a key to the plaintext, rune by rune, modulo 29. Call the plaintext runes p and the key k. Two ciphertext runes in a row come out equal exactly when

pi+1 − pi  =  ki − ki+1

The left side is a fact about English. The right side is a fact about the key. So the rate at which the ciphertext repeats itself is just how often English produces the particular letter-step that the key happens to be undoing at that moment. Whatever key you choose, it can never do better than the rarest step English has. Measure that on eight million runes of English put through the Gematria Primus, and the rarest step occurs 1.82% of the time.

Every additive cipher lives to the right of the wall, whatever its key, interrupters and all. The book is on the left.

The floor is generous twice over: to reach it the key would have to sit on that one rarest step at every single position, which makes it an arithmetic progression, which makes it periodic with period 29, which the period scan rules out. A key that is not doing that lands at 3.45%, where every machine below lands. And the wall is not an artefact of assuming English: the rarest step was measured in seven languages, and none of them has one below the book. Modern English 1.81%, Middle English 1.61%, Latin 1.26%, Italian 1.02%, Spanish 1.01%, German 1.00%, Dutch 1.00%, Old English 1.06%. The weakest case still leaves the book 3.8 standard deviations under the floor.

Run the machines yourself

Rather than take the argument, take the experiment. Each of these enciphers real English, written in the book's own runes and spelling conventions, and gets scored on the one statistic that carries information. The book's line is marked.

machinerepeats

The red line is chance, 1 / 29 = 3.448%. Every machine is seeded afresh each run, so the numbers wobble; none of them wobbles anywhere near the book.

Several of these match the book perfectly on flatness and on having no key length: a running key, a one-time pad, a rotor machine, Byrne's Chaocipher. None of them matches it on the one thing that is not noise.

Try to hit it

The last box is the honest one. Pick a cipher, give it any key you like, and see whether you can land on 0.664%. You cannot, and the reason you cannot is the inequality above.

 

Pick a cipher and press encipher.
The last option is the exception that proves the rule: alphabets picked on purpose to dodge English's own letter pairs can get under the book. That is the only class the measurement leaves standing, and here is how narrow it is. Draw an alphabet pairing at random and its repeat rate averages exactly 1/29. Draw four million of them and two reach the book's 0.664%; the best one in a thousand sits at 1.25%, the best in a million at 0.70%. No natural algebra of the Gematria Primus gets close either: Atbash 2.66%, multiply by thirteen 2.55%, cube 2.79%, inversion mod 29 2.02%, the best shift 1.82%. So the surviving cipher is not one somebody picked a key for. It is one somebody built already knowing what English does.

Can anything tell it from noise?

Every result above is a negative, and negatives pile up into a question worth asking straight out. Write down the smallest generator that could have produced these 12,956 runes. It needs three numbers:

12,956 runes long  ·  the observed rune frequencies  ·  one repeat-suppression factor

No plaintext. No key. No structure of any kind. Now run eleven statistics on the book and on sequences out of that generator, and see whether anything separates them.

Two of the eleven are fitted by the generator and so cannot fail; they are reported and thrown out of the verdict. And a battery too blunt to detect anything at all would make "indistinguishable" a meaningless word, so the same eleven statistics and the same matched-null procedure are run on ciphers whose answers are already known. Those controls are the actual result here. The book's own bar only means something next to them.

English, no cipher 110.6
Vigenere, key of 20 73.8
Vigenere, key of 300 68.2
running key 2.2
one-time pad 2.0
the Liber Primus 2.3
How far each sequence sits from a three-parameter message-free generator fitted to it, in standard deviations, on whichever of the nine unfitted statistics separates it best. The battery finds a 300-long Vigenere at 68 sigma, which the ordinary period scan walks straight past. It finds nothing in the book.

So the battery has teeth, and the book has nothing in it that three numbers do not already account for.

And here is what that does not mean. Look at the two bars above the book. A running key and a one-time pad also pass, at 2.2 and 2.0, and both of those unquestionably carry messages. Passing is not evidence of absence. The correct statement is narrower, and more useful than the overclaim would have been:

The unsolved Liber Primus sits in the same statistical class as a one-time pad. Whatever is in it is not reachable by statistics of this kind.

the period nobody scanned for

One cipher did fit the evidence, briefly, and it is worth showing because of how it died.

The suppression is not just present, it is featureless. The 86 surviving repeats favour no rune (chi-square 27.8 on 28 degrees of freedom), no page block, no word boundary, no position beside the book's own F interrupter, no period, and no part of the text. There is not one triple in the corpus. A flat, position-independent suppression like that is exactly what a polyalphabetic cipher gives when the relation between consecutive alphabets is fixed: each alphabet is the last one composed with a single permutation g. Then the repeat rate is g's bigram weight at every position, the same number everywhere, which is what the page shows.

That cipher is periodic, and its period is the order of g, which has nothing to do with the length of a key. For a permutation of 29 symbols that order runs as high as 2520. Every period scan in this study stopped at 150, on the reasonable-sounding grounds that keys are short.

Rescanned to 2600, by two instruments, against 200 matched nulls carrying the book's own length, marginals and suppression: the strongest coincidence lag reads 4.37% where a real period would read 6.1%, and the strongest coset index of coincidence reads 1.09 where a real period would read 1.78. The book's scan peaks at 4.2 sigma; the nulls peak at 3.8 on average and 4.7 at worst. A scan 6400 lags wide always finds something, and this is that. The last cipher whose signature matched is gone with it.

Why the book resists a word list

Everything above is about what the cipher cannot be. This is about why the obvious attack fails even where it ought to work, and it turns on a detail that had never been checked.

The book's two solved Vigenere pages leave a rune unenciphered and do not advance the key there. Every description of that rule, including this study's own earlier code, calls it "an F rune in the ciphertext". There are three candidate rules and they are separable by one question, because a Vigenere key has to come out exactly periodic:

rule0_welcomejpg107-167
skip where ciphertext is Fno periodno period
skip where plaintext is Fperiodic at 8 · DIUINITYperiodic at 13 · FIRFUMFERENFE
skip nothingno periodno period
Only the middle row works, and the two are not interchangeable. On 0_welcome there are 25 ciphertext F runes and only 11 interrupters, so fourteen are ordinary encipherments that happen to land on F. Every interrupter is a ciphertext F, so the rule is visible in principle. Most ciphertext F runes are not interrupters, so it is invisible in practice, and the ciphertext is all a solver has.

The cost is not subtle. Decrypt 0_welcome under the ciphertext-F rule using the true key and you get 68 of 515 runes right, 13.2%, with the first error at offset 5. The right key looks exactly like a wrong key.

And that is why a decade of dictionary attacks has failed. Running a 73,240-key dictionary against 0_welcome, a page whose key is the dictionary word DIVINITY, puts the true key at rank 5,294. The control fails, so every negative such a sweep produces is a statement about the interrupter and not about the book.

Worse: one of the book's two keys is not a word at all. DIUINITY is DIVINITY in the book's spelling. FIRFUMFERENFE is CIRCUMFERENCE with every C replaced by F, and searching 73,604 dictionary entries for anything whose rune spelling matches it returns nothing. A plain word list can reach one of the two keys the Liber Primus has actually demonstrated and structurally cannot reach the other.

the attack that does work

The interrupter is rare, 2.14% of runes on one page and 0.63% on the other, so a text runs clean before the first one. On both pages the first interrupter falls at offset 48 and 49. A prefix shorter than that is an ordinary Vigenere at phase zero, and there the attack finds the key:

16 runes · 0 intrrank 2
24 runes · 0 intrrank 1
32 runes · 0 intrrank 1
40 runes · 0 intrrank 1
48 runes · 0 intrrank 1
64 runes · 1 intrrank 1
96 runes · 3 intrrank 1
whole page · 11 intrrank 81
Bar length is the z-score of the best-scoring key; the label on the right is where the TRUE key ranked. At 48 runes the attack reads out WELCOMEWELCOMEPILGRIMTOTHEG, which is the page. The middle rows correct the obvious story: it is not the FIRST interrupter that breaks the attack, since 64 runes contains one and 96 contains three and both still rank the true key first. The score simply decays as they accumulate. It takes the whole page, with eleven, to lose the key, and there the attack returns a confident wrong answer at z +8.89.

So the instrument works. Sweeping every dictionary word and every single global rune substitution of one gives 13,257,833 candidate keys, and that version recovers both of the book's keys at rank 1: DIVINITY at z +3.98, and CIRCUMFERENCE with C replaced by F at z +7.48, reading out ACOANDURNGALESSONTHEMASTEREXPLAINEDTHE. Only once a search can find the keys the book actually used is its silence worth anything.

Run it on the first 44 runes of all nine unsolved blocks and nothing survives: the smallest empirical tail fraction anywhere is p = 0.32.

the largest crib, closed

The 51-rune red passage on p53 gets the strongest version: every dictionary word, at every key phase, under every interrupter hypothesis. The phase sweep matters because a crib in the middle of a block does not know where the key is; only the count of interrupters before it matters, not which ones, so the phase takes at most L values.

The interrupter hypothesis matters more, and the control shows why. Take a 51-rune window of 0_welcome that contains an interrupter and run the attack twice: with the hypothesis supplied the true key is found at z +9.06; without it, the same true key scores z −11.96. One unmodelled interrupter is the difference between solving the window and never seeing it. An attack that does not enumerate the hypothesis is testing its own luck.

The p53 passage holds exactly one ciphertext F, so there are two hypotheses, and both were run. No interrupter: best key HUMMINGBIRD at z +1.94. Interrupter at offset 38: best key ELSA at z +0.87. Against a true-positive band of +4.03 to +12.40 on the controls. The longest red passage in the Liber Primus is not a Vigenere under any of 73,193 dictionary rune-keys, at any phase, under either interrupter hypothesis.

The crib forces the key

And the book uses four ciphers, not three. Every account of the Liber Primus, this study's included, lists plaintext, Atbash, and Vigenere with a word key. p56_an_end is none of those. Its shift at key index j is (primej+1 − 1) mod 29, carrying the same plaintext-F interrupter, and it reproduces the page at 85 of 85 runes. For the first twenty-nine key indices that shift is identically the gematria value of the j-th rune minus one, which is the book's own table read off as a key.

That matters for everything else, because a running key is aperiodic and lies outside every period-based exclusion here. Tested separately against the same fixed stream at 2,000 starting offsets, with and without an interrupter: no crib of eight runes or more fits it at any offset. The control finds p56's own heading ANEND at offset 0.

The word-list attacks above search key space and score what comes back, which needs a null and yields a probability. There is a better move available, and it yields a count.

A crib's shape determines the key. If a word of length L sits at offset o in the span and the cipher is a periodic Vigenere of period p, then choosing that word fixes k[(o+i) mod p] = c[o+i] − w[i] for every i. So when p is at most L, one word determines the key completely; any internal inconsistency rejects it outright; and whatever survives must then decrypt the whole span into dictionary words. Iterating words therefore enumerates the entire family exactly, where brute-forcing key streams cannot. The deepest crib below stands for 366,451,025,462,807,220 key streams and clears in about a second.

Each crib is anchored on its longest word, which sets how far the exclusion reaches. A period equal to that length is degenerate: the anchor pins the key with nothing left over, so a hit is guaranteed whenever the remaining words happen to be words. Only periods strictly below the anchor carry information, and the table keeps them apart.

cribrunesanchorperiodskey streamssurvivors
p0-2 @01381..85.2×10110 (7 degenerate at 8)
p3-7 @016111..111.3×10133
p15-22 @91151..52.1×1070 (32 degenerate at 5)
p23-26 @01661..66.2×1080
p27-32 @019121..123.7×10170 (5 degenerate at 12)
p33-39 @01081..85.2×1011320
p54-55 @0981..85.2×10111,441
p40-53 @295751111..111.3×10130, at any period
Every periodic key each crib admits, enumerated exactly rather than sampled. The longest red passage in the book, 51 runes at the end of p53, admits none at all.

The control is what makes those zeros mean anything. The same code path ran on the twenty headings whose plaintext is published. Fifteen have a true period the anchor can reach, and the true plaintext comes back on 15 of 15, Atbash pages included.

That control earned its keep immediately. The first run recovered 11 of 15 and missed four, all Atbash. The cause was a sign error: Vigenere is c = p + k so k = c − p, Beaufort is c = k − p so k = c + p, and the code had k = p − c. Atbash is Beaufort with a constant key, so the bug lost exactly and only the Atbash pages, and nothing else in the run looked wrong. Four rows in a control table caught it.

Five of the ten informative cribs admit no periodic key at all below their anchor length, and the longest crib in the book admits none at any period whatsoever. The degenerate hits look degenerate: p27-32's five read UANPROCLIUITIESPRUT and GABCRYSTALLISEDCRAY, which is what a twelve-rune anchor pinning a twelve-long key leaves behind. Summed over the informative cribs, the family enumerated and cleared is 391,725,064,994,350,495 key streams, and it needed no null at all.

What this is and is not

It is not a solution. The Liber Primus is not open at the end of this page, and this page does not claim a key, a plaintext, or a method. What it claims is a shape for the search.

The doubled-rune deficit is not new. The solving community found it years ago and reports the same count of 86; it is the one clue everyone agrees on. What the Artificial Wasteland adds is the arithmetic that turns the clue into a bound, and the bench that shows sixteen cipher machines failing to clear it. The upshot is a negative result with a sharp edge: the additive family is excluded at 9.8 standard deviations, and the affine family at 6.3, against any English-like plaintext. That family is where the search has been.

Two readings survive, and both are narrower than they were. Either the book uses a cipher whose alphabets systematically avoid the plaintext's own letter pairs (a class one random alphabet pairing in two million reaches, so a deliberate signature rather than an accident of a key), or the unsolved pages carry no message and a generator with a no-repeat rule produced them. What has been taken away from the first reading since: shuffling the plaintext before enciphering does not help, and in fact raises the floor from 1.81% to 2.39%; no fixed linear window over the plaintext helps either; and the one cipher whose signature actually matched the flat, featureless suppression the book shows is excluded out to period 2600, which covers every order a permutation of 29 symbols can have. The text still cannot tell the two readings apart, and this page will not pretend otherwise, though the section below on how the pages were made leaves the two no longer evenly matched. What it can say is that the 86 repeats are really on the page: the pipeline behind this study re-read the book off its own scans by machine and confirmed 84 of the 86 directly, with zero disagreements against the human transcription anywhere in the unsolved pages.

It did find two errors, both on a solved page, and they are worth a sentence for what they do. One transcription reads four runes where the scans show three at two spots in the koan. That same four-rune form appears five times on the page and the machine reads four at three of them, so it is two isolated slips rather than a transcriber who cannot tell the glyphs apart, and they land on the koan's hinge. The page turns four times on the difference between who and what: that is not who you are, that is only what you are called; that is what you do, not who you are; that is only your species, not who you are; that is merely what you are, not who you are. In the published transcription the first two collapse into tautologies. The last two were already right, which is what shows the pattern.

Where this is weakest, since someone should say it. The bound is a property of the plaintext, and the plaintext used to measure it is 8.0 million runes of Project Gutenberg English. The Liber Primus's own plaintext is 2,977 runes, and it is measurably not that reference: its adjacent-difference distribution against English proportions gives chi-square 82.7 on 28 degrees of freedom, p below one in ten thousand. It is spikier, which is why its own floor sits at 1.31% rather than English's 1.81%. And 2,977 runes cannot pin a minimum: the 99.9% interval around that 1.31% reaches down to 0.637%, which does not exclude the book at all. The direct test still does, at p = 2.3e-3 once corrected for trying all 29 differences, but the strong exclusion above rests on the large reference corpus and not on the book's own voice. That is the load-bearing assumption, and this is where it would break.

the colour channel, and the section map

Everything above reads the rune stream. The page images hold a second channel: ink colour. It is bimodal to the pixel. Of the 13,867 glyphs on the numbered pages, 13,631 carry no red at all and 236 are red through, with nothing in between. And on every page whose plaintext is known, the red runs read A KOAN, AN INSTRUCTION, KNOW THIS, AN END, PARABLE, WELCOME, A WARNING, SOME WISDOM, THE LOSS OF DIVINITY, PRESERVATION, ADHERENCE. Red marks the section heading. Page 57, which is not enciphered at all, has exactly its first seven runes in red and they spell PARABLE; page 56, which is enciphered, has exactly its first five and they decrypt to AN END. A second signal agrees independently: headings sit inside a pair of chapter marks, so the punctuation alone brackets them.

Neither signal is enciphered, so the unsolved book can be sectioned: thirteen headings, at exact positions, with their word lengths given. Word shapes survive any rune-wise cipher, which makes these the tightest cribs anyone has, and two of the thirteen contain a one-rune word, which in English is A or I. And the map says one thing by itself: the book reuses its headings heavily in the part we can read, twenty bracketed units over fourteen shapes, and not one of the twelve distinct unsolved shapes matches any of them. Relabelling the same 33 headings at random, the expected overlap is 3.2 and the chance of seeing none is p = 0.008. Whatever the unsolved sections are called, they are not called what the rest of the book calls things.

And the longest red passage in the book is not in that map at all. The heading detector caps a heading at thirty-four runes, on the sound grounds that a heading is short, which is exactly why it misses this one. It is 51 runes on p53: the last fifty-one runes of the whole p40-53 block, four lines of it, starting mid-line immediately after a comma, terminated by the four-dot chapter mark, with the cicada emblem printed directly beneath it. The next longest red passage anywhere in the Liber Primus is twenty-nine runes. Checked back at the pixels rather than taken from the derived file: of the 183 glyph units on p53, 132 sit at a redness of exactly 0.0 and 51 sit between 171 and 180, with nothing in between.

What red means is settled by the pages we can read. Twenty-one red passages there have known plaintext; one is a single glyph, and the other twenty are, without exception, WELCOME, A WARNING, WISDOM, SOME WISDOM, KNOW THIS, A KOAN, A PARABLE, AN END, AN INSTRUCTION three times, PRESERVATION, CONSUMPTION, ADHERENCE, THE LOSS OF DIVINITY, FOR ALL IS SACRED, and AN INSTRUCTION / COMMAND YOUR OWN SELF. Red never marks ordinary body text on any page whose plaintext is known. But only one of those twenty is long enough to be a passage rather than a label, so "long red means an instruction" is a prior with n = 1, and this page will not call it a rule.

What is certain rather than likely is the shape, because word lengths survive every rune-wise cipher. (5, 4, 4, 11, 2, 3, 2, 6, 5, 5, 2, 2) is a plaintext fact about the unsolved book, and at twelve words it is the largest single crib in it, sitting in the most distinguished position the book has.

The crib against the book's own dictionary, and a line this page had stated misleadingly. An earlier version noted that a five-million-word English corpus holds no phrase matching that twelve-word shape, in a tone inviting you to find it surprising. It is not, and the arithmetic says so: multiplying each slot length's frequency across two million words of English through the Gematria Primus gives p = 9.8×10−12 per window, so the expected number of matches is 0.00. Finding none is exactly what chance predicts, and the absence carries no information whatever.

What the crib does constrain is vocabulary. Measuring word length in runes, the only length a rune-wise cipher preserves, the book's 2,977 known plaintext runes give 754 words, 300 of them distinct. Eleven-rune words are rare, and the book has exactly three: PRESERUATIAN, UNREASONABLE, ENLIGHTENED. Every other slot in the passage has between twenty-seven and sixty-seven candidates; the fourth has three. Not a solution, and the passage is free to use a word the sample never shows. But the fourth word of the longest red passage in the Liber Primus is eleven runes long, and everywhere the book can be read it uses only three words of that length.

pages 49 to 51, and four characters the transcription has wrong

Three of the numbered pages carry no runes in their body at all. They carry a grid of 256 two-character tokens over [0-9A-Za-z], in 32 rows of 8. This study had excluded them entirely, because the corpus keeps runes and these pages have none. The solving community reads that grid in bases 59 to 64 and reports a standing problem with it: transcription disputes, capital against lowercase above all.

That dispute is decidable. The characters a reader confuses in this typeface are featureless vertical bars, and what separates them is height, and the grid supplies its own ruler: x-height 44 px (a c e m n o r s u v w x), cap height 67 px (B D E F H K L P T X Y Z and every digit, 1 included), ascender 75 px (b d k). Three clean bands, nothing between. So a bare bar at 67 px is a capital I and one at 74 px is a lowercase l.

The pages hold 13 bare bars. Three carry a dot, and those three are exactly the ones transcribed i, j, j. The other ten are undotted and 7 or 8 px wide, where a real L is 26 px wide with a foot and a real 1 is 19 px wide with a flag. Six of the ten confirm the transcription. Four contradict it: p49 row 5 1L and row 6 0L are both 74 px, so 1l and 0l; p50 row 10 0l and p51 row 7 3i are both 67 px with no dot, so 0I and 3I. Under any base-62 reading those four tokens change value, because L is 21 and l is 47, i is 44 and I is 18. Four of the 256 values in the block are wrong in the published transcription, in exactly the class the community already suspects.

A near miss, because the discipline is the point: the first run of this measurement reported six corrections, including both j's. It was wrong. The component filter had a minimum height, so a dot was never in the candidate list at all, and a dotted j was indistinguishable from a capital I by construction. The measurement said one thing and the pixels said another, and the pixels were right.

With the grid corrected, the base is decidable too, and seventeen tokens settle it. Order the alphabet 0-9 A-Z a-z. Some token's second character reaches x, index 59, so the base is at least 60. Now take the tokens whose first character is 4, the largest first character anywhere in the block: under base B those have value 4B plus the second digit, so they are the only tokens a byte ceiling constrains. There are seventeen, and every one has a second digit at index 15 or below. Base 60 forces exactly that, because 4 times 60 plus 15 is 255. Base 61 or 62 forces nothing, and seventeen tokens landing that low by chance is about one in ten billion. So: base 60, alphabet 0-9 A-Z a-x, values running exactly 0 to 255.

And then the bytes are nothing. 256 bytes is 2048 bits, exactly one prime factor of a 4096-bit RSA key, and Cicada's own signing key is 4096-bit; the grid is 32 rows of 8, so its columns give eight 32-byte chunks, the shape of eight SHA-256 digests. Fourteen readings of the grid, four block sizes, both endiannesses, 420 primality tests and not one prime. No chunk repeats at any size. Entropy 7.170 bits per byte against 7.18 for uniform random, and 161 distinct values against 161.9 expected. The block is indistinguishable from 256 random bytes. The correction is still worth having, because every published attempt on that block was run against four wrong values, but it does not make the block say anything.

the leakage, and what it is not

A rule that forbids repeats gives zero. The book gives 86, spread as a clean Poisson process at 19.2% of the chance rate. That gap is the sharpest handle anyone has, and two ordinary explanations for it are now dead. Not typesetting errors: the solved pages each decrypt under a rigid rule, so a typo would show as a position where the rule breaks, and they break at 0 of 2,963 positions, which caps the book's error rate at 0.10% where the leakage would need 0.66%. Not the seams between runs of a slowly-changing alphabet: fit that model to the repeat rate and it predicts a lag-2 rate of 3.95% to 5.87%, where the book's is 3.404%, which is 1/29 exactly. For every language as written the model cannot even reach the book's repeat rate. So the alphabet changes at essentially every position, and the whole weight falls back on those one-in-two-million pairings. Or the plaintext is not a language.

two things no cipher can change

Word lengths and punctuation are not enciphered. They pass through any rune-wise substitution untouched, which the solved pages confirm: the solved ciphertext and its own plaintext have the same word-length distribution to p = 0.798. So anything odd about them is a fact about the text underneath. Two things are odd. The unsolved pages are short of two-rune words: 15.5% against 24.1% in the book's own plaintext (p = 6×10−8) and against a range of 20.7% to 24.8% across fourteen English texts. It holds uniformly across all nine unsolved page groups. And they barely punctuate: 73.6 runes to a sentence against 29.2 in the book's own prose, and 0.31 commas per thousand runes against 12.43. Neither is proof of anything alone, since a different register could account for long unpunctuated passages. What they are is a test no proposed solution has had to pass: a correct decryption must come out 15.5% two-rune words, 74 runes to a sentence, and almost never using a comma.

how the pages themselves were made

The rune statistics cannot settle one thing that matters, and the image files can. Every JPEG carries its encoder's fingerprint in its quantisation tables and its embedded colour profile. Across the 58 numbered pages as served from Cicada's own onion service, all 58 share one quantisation table and one ICC profile, and that profile reads "Artifex Software sRGB ICC Profile", Copyright Artifex Software 2011. Artifex is Ghostscript, so the book was rendered from a document rather than scanned, which is also why its runes template-match to thirteen pixels of separation. None of the 58 carries appended data. The seventeen unnumbered pages carry none of that profile and use three other tables: they were not part of that render.

So pages 0 to 57 came out of one document in one act, and that batch holds both the fifty-six unsolved pages and the two solved ones, page 56 and page 57. Whatever made the unsolved pages made two pages that demonstrably say something, at the same time, from the same file. It is not proof that the rest says anything, since one document can hold filler and end-matter together. It does close off the easiest version of the no-message reading, the one where the unsolved pages were padding bolted on to a real book.

a lead that looked real and was not

The strongest key lengths in the period scan were 15, 30, 45, 60, 120 and 345, all multiples of 15, and the pattern replicated independently in both halves of the book. It is an artefact. Coset scores at multiples of a period are statistically dependent, so one fluctuation at 15 shows up six times and looks like corroboration. Against message-free controls carrying the book's exact length, marginals and repeat-suppression, each control taking its own best divisor, the book's score is beaten by 9 of 120 controls. It is recorded here so the next person does not spend a week on it.

the check

Every number on this page is recomputed from the committed corpus by research/cicada-3301/verify.mjs: the Gematria Primus, the corpus sizes, all nine solved pages re-derived from raw ciphertext, the flatness tests, the 86, the lag and period scans, the three exclusion bounds by explicit minimisation (29 shifts, 812 affine maps, and a Hungarian assignment over the bigram matrix), ten cipher machines, and the OCR audit. It exits non-zero if any of it moves. The study, the traps, and the file list are in research/cicada-3301/README.md.

Sources: page scans and one transcription, relikd/LiberPrayground; two more, rtkd/idkfa and Taiiwo/cicada, which agree with it exactly on all 12,956 unsolved runes. English reference: fourteen Project Gutenberg texts, 8.0M runes through the Gematria Primus.