A census · pre-registered · 29 September 2026
Ink With Nothing Under It
When you search inside a scanned book at the Internet Archive, you are searching a text layer that a machine made from the page images. Usually it is garbled here and there. Sometimes it is not garbled at all, because it is not there: a printed line on the page has no text under it, so no search will ever land on it. This page counts how often that happens on a random page of a random pre-1929 book or newspaper, and shows every case it found.
Three lines of Blondel
This is page 675 of François Blondel's Cours d'architecture (Paris, 1683), from the copy the Internet Archive scanned in 2010 (item coursdarchitectu00blon, page n894). Press the second button. Ink that sits under one of the text layer's own word boxes turns blue. Ink that no word box covers turns red.
detect.py, from the item's own hOCR for this leaf (leaf 897 of the scan).The three red lines, as printed (the long s kept):

Elles ne doivent pourtant point avoir moins de quatre pouces,
leur largeur jamais moins d'un pied ny plus d'un pied & demy. Les
Anciens faiſoient ordinairement leurs marches en nombre impair;
This is Blondel reporting Palladio's lower limits for a stair (a step no less than four inches high, its tread no less than a foot deep and no more than a foot and a half) and then the ancients' habit of building their flights in odd numbers. The text layer goes from marches contniucs (its reading of continües) to Il n'en faut jamais. Of the three lines between, it holds one thing: a three-character fragment, im[, with zero confidence, at the place where impair; is printed. There is no other version of them, mangled or not, anywhere in the book's text. Ask the Archive's own search inside this book for "doivent pourtant", "moins de quatre pouces" or "ordinairement leurs marches" and it finds nothing, while a word from the same page that is covered, Chambor, is found on it (the probes).
This ground found the gap by accident, building a page on the stair rule from the page images rather than the text. The question it left was how often it happens. Nobody seemed to have counted, so this counts.
The census
Before anything was measured, the method, the sample size, the rules for excluding a page, the categories and four guesses were written down and committed (the pre-registration; this project's repository has it at commit 69c8f6f6, before the draw at de9c9dd5). What changed along the way is in the deviations. Then:
- Draw. 300 items from the Archive's own search for scanned texts with OCR: 100 dated before 1800, 100 from the 1800s, 100 from 1900 to 1928. The search's random sort takes no seed, so the drawn list was committed and is the sample. One page from each, chosen by a seeded hash of the item's name.
- Fetch the page's image and the item's own hOCR for that page (the file its full text and search are made from). Check that they are the same page: where the image was found by page number rather than by the file name the hOCR records, a separate OCR engine reads five of the hOCR's lines off the image and must agree with them. Five pages failed that and were dropped; one of the pilot's had been off by exactly one page.
- Look for uncovered ink. Colour the ink the text layer's boxes cover, and flag any band of ink the height of a line that runs across a stretch of the text column with nothing over it. This step runs no OCR of its own and decides nothing. It only says where to look.
- Look. All 656 flagged bands were shown, as the crops below are, to eight classifiers (AI agents, each given the same brief) that were told the categories and not the question, and then every call of dropped text and a random tenth of the rest were looked at again by me.
Of the 300 items, 213 gave a usable page. The rest: 68 had no hOCR at all (all 68 were older scans whose OCR the metadata credits to ABBYY FineReader), 8 files would not download, 5 failed the same-page check, 4 were access-restricted and 2 held several books in one item.
| Pages | Usable | Missing body text | Share (95% interval) | Missing any printed text |
|---|---|---|---|---|
| Printed before 1800 | 70 | 8 | 11.4% (5.9 to 21.0) | 19 |
| 1800 to 1899 | 66 | 9 | 13.6% (7.3 to 23.9) | 17 |
| 1900 to 1928 | 77 | 9 | 11.7% (6.3 to 20.7) | 22 |
| Newspapers and periodicals | 103 | 19 | 18.4% (12.1 to 27.0) | 42 |
| Everything else | 110 | 7 | 6.4% (3.1 to 12.6) | 16 |
| All | 213 | 26 | 12.2% (8.5 to 17.3) | 58 |
The split between newspapers and everything else is a rough one, made by matching words in each item's title, collection and name, and a few items will be on the wrong side of it. The difference between the centuries is nothing the sample can tell apart. The difference between newspapers and everything else is the largest in the table: a page of a newspaper was about three times as likely to have lost body text. But the pre-registration named this split as descriptive only, and the word list that makes it was widened after the list of pages had been read, so it is a finding to test on a new sample, not a result.
Every page it found
Each card is the band the census flagged, twice: the scan, then the same scan tinted, red where there is ink and no text. The link opens the page at the Internet Archive, where you can compare the page with its full text yourself.
Loading the pages… (the list is in data.json)
Search for the missing words yourself
Each of these words or phrases is printed on the sampled page, in text the census found missing from the layer. The Archive's own search inside was asked for it on 29 September 2026, together with a control word that is printed on the same page in text the layer does cover. The control must be found on that page, or the probe proves nothing.
- Blondel, Cours d'architecture, 1683:
"doivent pourtant","moins de quatre pouces","ordinairement leurs marches": no hits anywhere in the book. ControlChambor: found on this page. - Boston American, 18 April 1910:
"Thinking Machine",Scituate,Hingham: no hits. ControlFutrelle: found on this page. - American Journal of Mining, 6 July 1867:
schlich: no hits. Controlstamped: found on this page. - The New York Saturday Press, 22 January 1859:
Bibliographical: no hits. ControlAddison: found on this page. - La Tribuna, 1918:
Reuter: no hits. ControlHindenburg: found on this page. - The Bombay Chronicle, 28 August 1928:
Kellogg,Lucknow: no hits. No control is possible: this page's text layer is empty.
The Boston American one is worth opening. The lines with no text under them tell how "The Thinking Machine," otherwise Jacques Futrelle, the analyst of crime, was driving home to Scituate on a Sunday and was held up in a police speed trap in Hingham. Futrelle wrote the detective stories of Professor Van Dusen, "The Thinking Machine," and died on the Titanic two years later. Neither his detective's name nor the town he was caught in can be found by searching that page.
How much it misses
The counts above are lower bounds, and the census measured by how much. On every usable page, up to three lines that the text layer does have were deleted from it on purpose, one at a time, and the detector was run again to see whether it flagged the hole.
| Width of the deleted line | Holes | Found | Recall |
|---|---|---|---|
| under a sixth of the text block | 145 | 43 | 29.7% |
| a sixth to a half | 206 | 142 | 68.9% |
| half or more | 274 | 223 | 81.4% |
| all | 625 | 408 | 65.3% |
A full line missing from a book page is found four times in five. The short last line of a paragraph, or a narrow newspaper column on a wide page, mostly is not. So the true share of pages with a line missing is higher than 12.2%, by an amount this census cannot pin down: recall is per line, and a page with several missing lines is likelier to be caught than a page with one.
What the look found, and how sure it is
The 656 flagged bands, as finally called: 86 body text missing; 188 other printed text missing (mastheads, headlines, advertisements, table rows, page numbers); 352 not text (rules, borders, pictures, page edges, stains); 13 covered after all (fragments of letters whose words are in the layer); and 17 handwriting. The last class was not in the pre-registration. Three sampled items are manuscripts (a Karnataka state archive file, a Baptist meeting's minute book of 1699 to 1708, and a page whose only uncovered ink is a handwritten mark in the margin), and the pre-registration defined a missing line as printed text, so their bands are counted as none of the other four.
Of the 333 bands I looked at again, I agreed with the classifier on 298 (89.5%). Seventeen of the other 35 were the handwriting. The rest were mostly calls of body text that I moved to other text, such as advertisement copy, which the brief assigns there.
A second, independent witness: a different OCR engine (tesseract 5.3.4, run here) read the 86 missing-body-text bands off the images. It read 1,150 words of five letters or more, and 917 of them appear nowhere in that page's whole text layer, not merely nowhere under the band. That rules out the obvious worry, that the words are in the layer and only their boxes are drawn in the wrong place. On 23 of the 26 pages it found such words; on the other three (black-letter type, the words under a printed music score, and a single line in a small-town bulletin) it read no word of five letters from the band.
The four guesses
| Written before looking | What happened |
|---|---|
| G1. Between 2% and 8% of pages lose body text. | Wrong. 12.2%, and its interval (8.5% to 17.3%) sits above the guess. |
| G2. Older scans OCR'd with ABBYY FineReader lose more than newer ones OCR'd with tesseract. | Not supported, and barely testable. 0 of 20 ABBYY pages against 26 of 193 tesseract pages; 68 of the ABBYY-era items had no hOCR to test at all. In practice this census measures the tesseract-era text layers. |
| G3. Most of what the detector flags that is not text will be rules, borders and show-through, and not-text will outnumber missing text. | Right on the count, 352 against 274. The first half was never scored band by band; a keyword count over the classifiers' notes finds a rule, border, box, frame, edge, margin, gutter, background or show-through named in 270 of the 352. |
| G4. Recall above 80% for lines wider than half the block, below 50% for lines narrower than a sixth. | Right, narrowly on the first: 81.4% and 29.7%. |
The check
Every figure on this page is recomputed from the census's own per-page records by verify-ink-with-nothing-under-it.mjs, which also finds each figure in this page's text and goes red if a record is changed. With --live it goes to archive.org itself, fetches the hOCR for the Blondel page, and confirms that no word box lies over the three lines, that none of their words is in that leaf's text, and that Chambor is. From an empty directory:
curl -sO https://artwaste.land/checks/verify-ink-with-nothing-under-it.mjs && node verify-ink-with-nothing-under-it.mjs --live
The census's records are served beside this page: census.json (every drawn item, its status, its counts and its planted-line results) and data.json (the 26 pages). The programs that made them, the sample, the adjudication and every classifier's call are in research/where-the-ocr-drops-a-line/; the detector is detect.py.
What this does not say
- It is not a measure of how accurate the text is. A line read as nonsense counts here as covered. This counts only absence.
- It is a snapshot. The Archive re-runs OCR on old items from time to time; these are the text layers it served on 29 September 2026.
- "Random" means the Archive's random sort over its search index, not a probability sample of printed pages. The population is what the Archive holds under these three date queries, and it leans heavily on microfilmed newspapers and library scans.
- The detector only looks inside the area where the text layer found some text. A drop at the very edge of that area, or a page where the layer found text only in one corner, is under-counted. Two printed pages have an empty text layer: a page of The Bombay Chronicle and the broadside The Parallel (1682). They count as pages with missing body text, but are left out of the per-line rate, because the layer has no lines to count against. (Two other sampled pages with empty layers are blank.)
- With 213 pages, the intervals are wide. No comparison between kinds of page was registered as a test; the newspaper difference is the one worth testing next.
- It says nothing about why. The missing bands cluster in dense newspaper columns, text beside pictures, black-letter and blurred type, and page edges, which is where layout analysis usually struggles, but the census was not built to test a cause.
Sources
- Every page image, hOCR file, scan-data file and metadata record was fetched from archive.org on 29 September 2026; each exhibit links to its page. Items named here:
coursdarchitectu00blon,bostonamerican19100418,india.history.resource.6852,sim_engineering-and-mining-journal_1867-07-06_4_1,per_new-york-saturday-press_the-new-york-saturday-press_1859-01-22_2_4,IEI0105530_1918_00241. - The OCR engine for each item is the one its own metadata names (the
ocrfield). - Jacques Futrelle: Wikipedia, "Jacques Futrelle".
- The stair-rule page that found the Blondel gap: One Inch Up, Two Across.