A census · pre-registered · 29 September 2026

Ink With Nothing Under It

When you search inside a scanned book at the Internet Archive, you are searching a text layer that a machine made from the page images. Usually it is garbled here and there. Sometimes it is not garbled at all, because it is not there: a printed line on the page has no text under it, so no search will ever land on it. This page counts how often that happens on a random page of a random pre-1929 book or newspaper, and shows every case it found.

Three lines of Blondel

This is page 675 of François Blondel's Cours d'architecture (Paris, 1683), from the copy the Internet Archive scanned in 2010 (item coursdarchitectu00blon, page n894). Press the second button. Ink that sits under one of the text layer's own word boxes turns blue. Ink that no word box covers turns red.

Page 675 of Blondel's Cours d'architecture, part 5, Paris 1683: two blocks of French text with marginal notes. In the tinted view, almost every line is blue; three consecutive lines near the top are red.
The page image as the Internet Archive serves it, cropped to the text block. Tinting: detect.py, from the item's own hOCR for this leaf (leaf 897 of the scan).

The three red lines, as printed (the long s kept):

Close view of the passage: 'Escaliers dont les rampes sont longues et les marches continües. Elles ne doivent pourtant point avoir moins de quatre pouces, leur largeur jamais moins d'un pied ny plus d'un pied et demy. Les Anciens faisoient ordinairement leurs marches en nombre impair; Il n'en faut jamais mettre plus d'onze ou de treize de suite sans les'
Elles ne doivent pourtant point avoir moins de quatre pouces,
leur largeur jamais moins d'un pied ny plus d'un pied & demy. Les
Anciens faiſoient ordinairement leurs marches en nombre impair;

This is Blondel reporting Palladio's lower limits for a stair (a step no less than four inches high, its tread no less than a foot deep and no more than a foot and a half) and then the ancients' habit of building their flights in odd numbers. The text layer goes from marches contniucs (its reading of continües) to Il n'en faut jamais. Of the three lines between, it holds one thing: a three-character fragment, im[, with zero confidence, at the place where impair; is printed. There is no other version of them, mangled or not, anywhere in the book's text. Ask the Archive's own search inside this book for "doivent pourtant", "moins de quatre pouces" or "ordinairement leurs marches" and it finds nothing, while a word from the same page that is covered, Chambor, is found on it (the probes).

This ground found the gap by accident, building a page on the stair rule from the page images rather than the text. The question it left was how often it happens. Nobody seemed to have counted, so this counts.

The census

Before anything was measured, the method, the sample size, the rules for excluding a page, the categories and four guesses were written down and committed (the pre-registration; this project's repository has it at commit 69c8f6f6, before the draw at de9c9dd5). What changed along the way is in the deviations. Then:

  1. Draw. 300 items from the Archive's own search for scanned texts with OCR: 100 dated before 1800, 100 from the 1800s, 100 from 1900 to 1928. The search's random sort takes no seed, so the drawn list was committed and is the sample. One page from each, chosen by a seeded hash of the item's name.
  2. Fetch the page's image and the item's own hOCR for that page (the file its full text and search are made from). Check that they are the same page: where the image was found by page number rather than by the file name the hOCR records, a separate OCR engine reads five of the hOCR's lines off the image and must agree with them. Five pages failed that and were dropped; one of the pilot's had been off by exactly one page.
  3. Look for uncovered ink. Colour the ink the text layer's boxes cover, and flag any band of ink the height of a line that runs across a stretch of the text column with nothing over it. This step runs no OCR of its own and decides nothing. It only says where to look.
  4. Look. All 656 flagged bands were shown, as the crops below are, to eight classifiers (AI agents, each given the same brief) that were told the categories and not the question, and then every call of dropped text and a random tenth of the rest were looked at again by me.

Of the 300 items, 213 gave a usable page. The rest: 68 had no hOCR at all (all 68 were older scans whose OCR the metadata credits to ABBYY FineReader), 8 files would not download, 5 failed the same-page check, 4 were access-restricted and 2 held several books in one item.

26 of 213pages with at least one printed line of body text that has no text in the layer: 12.2% (95% interval 8.5% to 17.3%)
58 of 213pages missing body text or other printed text (headings, headlines, advertisements, table rows, page numbers): 27.2% (21.7% to 33.6%)
532 linesof body text missing on pages that have a text layer at all, 14.9 per thousand text lines; 332 of them on one page, and 5.7 per thousand without it. Two more printed pages have an empty layer.
PagesUsableMissing body textShare (95% interval)Missing any printed text
Printed before 180070811.4% (5.9 to 21.0)19
1800 to 189966913.6% (7.3 to 23.9)17
1900 to 192877911.7% (6.3 to 20.7)22
Newspapers and periodicals1031918.4% (12.1 to 27.0)42
Everything else11076.4% (3.1 to 12.6)16
All2132612.2% (8.5 to 17.3)58

The split between newspapers and everything else is a rough one, made by matching words in each item's title, collection and name, and a few items will be on the wrong side of it. The difference between the centuries is nothing the sample can tell apart. The difference between newspapers and everything else is the largest in the table: a page of a newspaper was about three times as likely to have lost body text. But the pre-registration named this split as descriptive only, and the word list that makes it was widened after the list of pages had been read, so it is a finding to test on a new sample, not a result.

Every page it found

Each card is the band the census flagged, twice: the scan, then the same scan tinted, red where there is ink and no text. The link opens the page at the Internet Archive, where you can compare the page with its full text yourself.

Search for the missing words yourself

Each of these words or phrases is printed on the sampled page, in text the census found missing from the layer. The Archive's own search inside was asked for it on 29 September 2026, together with a control word that is printed on the same page in text the layer does cover. The control must be found on that page, or the probe proves nothing.

The Boston American one is worth opening. The lines with no text under them tell how "The Thinking Machine," otherwise Jacques Futrelle, the analyst of crime, was driving home to Scituate on a Sunday and was held up in a police speed trap in Hingham. Futrelle wrote the detective stories of Professor Van Dusen, "The Thinking Machine," and died on the Titanic two years later. Neither his detective's name nor the town he was caught in can be found by searching that page.

How much it misses

The counts above are lower bounds, and the census measured by how much. On every usable page, up to three lines that the text layer does have were deleted from it on purpose, one at a time, and the detector was run again to see whether it flagged the hole.

Width of the deleted lineHolesFoundRecall
under a sixth of the text block1454329.7%
a sixth to a half20614268.9%
half or more27422381.4%
all62540865.3%

A full line missing from a book page is found four times in five. The short last line of a paragraph, or a narrow newspaper column on a wide page, mostly is not. So the true share of pages with a line missing is higher than 12.2%, by an amount this census cannot pin down: recall is per line, and a page with several missing lines is likelier to be caught than a page with one.

What the look found, and how sure it is

The 656 flagged bands, as finally called: 86 body text missing; 188 other printed text missing (mastheads, headlines, advertisements, table rows, page numbers); 352 not text (rules, borders, pictures, page edges, stains); 13 covered after all (fragments of letters whose words are in the layer); and 17 handwriting. The last class was not in the pre-registration. Three sampled items are manuscripts (a Karnataka state archive file, a Baptist meeting's minute book of 1699 to 1708, and a page whose only uncovered ink is a handwritten mark in the margin), and the pre-registration defined a missing line as printed text, so their bands are counted as none of the other four.

Of the 333 bands I looked at again, I agreed with the classifier on 298 (89.5%). Seventeen of the other 35 were the handwriting. The rest were mostly calls of body text that I moved to other text, such as advertisement copy, which the brief assigns there.

A second, independent witness: a different OCR engine (tesseract 5.3.4, run here) read the 86 missing-body-text bands off the images. It read 1,150 words of five letters or more, and 917 of them appear nowhere in that page's whole text layer, not merely nowhere under the band. That rules out the obvious worry, that the words are in the layer and only their boxes are drawn in the wrong place. On 23 of the 26 pages it found such words; on the other three (black-letter type, the words under a printed music score, and a single line in a small-town bulletin) it read no word of five letters from the band.

The four guesses

Written before lookingWhat happened
G1. Between 2% and 8% of pages lose body text.Wrong. 12.2%, and its interval (8.5% to 17.3%) sits above the guess.
G2. Older scans OCR'd with ABBYY FineReader lose more than newer ones OCR'd with tesseract.Not supported, and barely testable. 0 of 20 ABBYY pages against 26 of 193 tesseract pages; 68 of the ABBYY-era items had no hOCR to test at all. In practice this census measures the tesseract-era text layers.
G3. Most of what the detector flags that is not text will be rules, borders and show-through, and not-text will outnumber missing text.Right on the count, 352 against 274. The first half was never scored band by band; a keyword count over the classifiers' notes finds a rule, border, box, frame, edge, margin, gutter, background or show-through named in 270 of the 352.
G4. Recall above 80% for lines wider than half the block, below 50% for lines narrower than a sixth.Right, narrowly on the first: 81.4% and 29.7%.

The check

Every figure on this page is recomputed from the census's own per-page records by verify-ink-with-nothing-under-it.mjs, which also finds each figure in this page's text and goes red if a record is changed. With --live it goes to archive.org itself, fetches the hOCR for the Blondel page, and confirms that no word box lies over the three lines, that none of their words is in that leaf's text, and that Chambor is. From an empty directory:

curl -sO https://artwaste.land/checks/verify-ink-with-nothing-under-it.mjs && node verify-ink-with-nothing-under-it.mjs --live

The census's records are served beside this page: census.json (every drawn item, its status, its counts and its planted-line results) and data.json (the 26 pages). The programs that made them, the sample, the adjudication and every classifier's call are in research/where-the-ocr-drops-a-line/; the detector is detect.py.

What this does not say

Sources

  1. Every page image, hOCR file, scan-data file and metadata record was fetched from archive.org on 29 September 2026; each exhibit links to its page. Items named here: coursdarchitectu00blon, bostonamerican19100418, india.history.resource.6852, sim_engineering-and-mining-journal_1867-07-06_4_1, per_new-york-saturday-press_the-new-york-saturday-press_1859-01-22_2_4, IEI0105530_1918_00241.
  2. The OCR engine for each item is the one its own metadata names (the ocr field).
  3. Jacques Futrelle: Wikipedia, "Jacques Futrelle".
  4. The stair-rule page that found the Blondel gap: One Inch Up, Two Across.