The Page the Text Layer Lost
The Uniform Type Committee's 1915 report shipped its decisive per-character legibility table on one landscape sheet inside a portrait book, and every scan of the volume put the sheet upside-down. Tesseract read its headers as 'SHATVA AOVUNOOV' and 'SHOIVA GINIL', which is 'ACCURACY VALUES' and 'TIME VALUES' rotated a hundred and eighty degrees, and everyone who searched the text layer for the numbers behind the report's rankings saw a wall of gibberish and moved on. Rotating the archive.org page image ninety degrees clockwise, re-running the OCR, and hand-transcribing each of the 334 cells against the printed 'Average' row for a checksum brings the table back: 84 rows, 70 punctographic characters, seventy-two of them with a single estimate and thirteen with a 'preferred + second' pair for the alternative-context condition the report's own note describes; the character 1-4 gets a third estimate for accuracy only. The bottom-right 'Average' the report prints as 153.01, 123.89, 92.56, 96.08 does not exactly reproduce as the arithmetic mean of any obvious subset of the recovered values (the closest rule, one row per character preferring the 'preferred' estimate, gives 154.83, 124.45, 92.00, 95.42); the residual is small in three columns and larger in the Part-Word time, and the 1915 arithmetic is not reachable from the printed numbers alone.
· New York Point · braille · tactile writing · Uniform Type Committee · 1915 · legibility · punctographic characters · primary sources · OCR · text layer · Internet Archive · record-correcting · recovery · show-the-check · reproducibility · history of technology · accessibility
A century of readers searching the archive’s text layer for the numbers under the Uniform Type Committee’s 1915 verdict against New York Point saw the string SHATVA AOVUNOOV and moved on. It is ACCURACY VALUES, rotated. The whole table is under it.
The page
The Uniform Type Committee’s Fifth Biennial Report (New York, 1915) is the document that ended the War of the Dots. Its verdict against New York Point rested on rankings tabulated across five tactile alphabets, and the rankings themselves rested on Appendix B, which prints, for each of seventy punctographic characters, four numbers: a Part-Word time value, a Whole-Word time value, a Part-Word accuracy percentage, and a Whole-Word accuracy percentage. Anybody who wanted to check the arithmetic under the rankings had to look at Appendix B.
The report is on the Internet Archive, freely readable, at fifthbiennialrep0000unse. The full text layer, run through tesseract by the archive’s own pipeline, is 24,367 lines of OCR. Anybody who searched inside the report through the archive would find, at the position where Appendix B ought to be, a block of nearly five hundred lines like this:
-$0°96
92°86
00° 001
9¢" 26
19°¢6
...
SHATVA AOVUNOOV
...
SHOIVA GINIL
The block is unreadable. It sits between the end of the report body on page 45 and the start of Appendix C on page 47, and it accounts for the missing page 46 in the archive’s page-number index. The whole appendix is inside it.
What the OCR was looking at
Appendix B is a wide table: seven columns and about forty-five rows, twice, sitting side by side to fit on one sheet. To make it fit inside a portrait-format book, the printer set the sheet sideways. Every other page of the volume reads bottom-of-the-page-toward-the-spine; this one reads bottom-of-the-page-away-from-the-spine, so that a reader can hold the book on its side to read the wide table. The archive’s scanner recorded the sheet without knowing that, and the archive’s OCR pipeline read the resulting image at the same orientation as the rest of the book, which put the table upside down.
An upside-down word, read left to right, is the original word reversed and each letter turned a hundred and eighty degrees. Rotated E looks like H; rotated L looks like T; rotated A looks like V; rotated U looks like A. ACCURACY VALUES reversed is SEULAV YCARUCCA, and applying the letter rotations gives SHATVA AOVUNOOV, which is what the OCR wrote. TIME VALUES under the same rule comes out as SHOIVA GINIL. The digits behave similarly: rotated 6 becomes 9, rotated 5 becomes something a lot like S, and the decimal point sits at the top of the character where a degree sign lives, so 96.05 reversed and rotated reads to tesseract as -$0°96.
Once the transformation is named, decoding a few of the values by hand becomes tractable, and enough of them cross-check that the shape of the table is unambiguous. But the digits 2 and 5 degrade under rotation, and so do the letter combinations in the character labels; a decoded reconstruction from the archive’s OCR alone would be unreliable, and would not carry the internal check the table itself supplies.
The fix
The two commands that recover Appendix B are:
curl -sSL https://archive.org/download/fifthbiennialrep0000unse/page/n47_w2000.jpg \
-o leaf-47.jpg
convert leaf-47.jpg -rotate 90 leaf-47-r90.jpg
tesseract leaf-47-r90.jpg out -l eng --psm 6
The first grabs the archive’s derived JPEG of the leaf that carries the appendix. The second turns the image ninety degrees clockwise so the table sits upright. The third re-OCRs the upright image with tesseract’s default English model. The tesseract output is a clean readable table; the archive’s is not.
The recovery in this project’s own repository is under research/tactile-codes/appendix-b/. Its raw materials are the page image and the tesseract TSV of it; its committed artifact is data/table.json, holding all 84 rows; its checker is verify.mjs, which runs three consistency invariants and one cross-check against the printed ‘Average’ row on the same page.
The table
There are 70 unique punctographic characters in Appendix B, labelled as hyphen-separated lists of the raised-dot positions that form them (the New York Point system numbers cells 1 through 8 by their vertical column pair, and the report writes 1-2-4-5 for the sign whose dots sit at positions 1, 2, 4 and 5). Thirteen of those characters carry two estimates, preferred and second, for the two alternative reading conditions the note at the top of the table describes; the character 1-4 gets a third estimate, third, for accuracy only. That gives 84 rows in total, and 334 non-null numeric values.
| character | Part-Word time | Whole-Word time | Part-Word accuracy | Whole-Word accuracy | qualifier |
|---|
What is inside the table
The rankings the report published in the same volume take these values as their input, and running the values through even the simplest sorts already surfaces the shape of the committee’s argument, and one or two figures the report did not quite state.
The slowest signs for Part-Word reading, at the top of Appendix B. Four characters tie at 213.68 or 213.70 (six-dot compound signs occupying two two-dot cells): 1-2-3-4-5, 1-2-3-5-6, 1-2-4-5-6, 2-3-4-5-6. These are the compound signs the reader has to disambiguate against nearly every other compound sign of similar shape, and the times are more than twice those of the single-dot 1 (92.38) and of the two-dot 1-5 (100.20). The 1-6 sign, one dot in one cell and one in the other, comes in at 189.93 and is on its own as the slowest non-compound.
The least accurate signs. The 1-6 character again, at 72.02 per cent Part-Word accuracy, is at the bottom. Two other signs are inside 80 per cent: 3-4 at 77.65 (a full column of dots in the right half of one cell, with nothing in the left half) and 1-4-7 at 80.36 (a three-dot compound whose second cell has a solitary top-of-column dot). These are the signs the committee’s argument against New York Point rests on: not the average sign, but the signs at the bottom of the accuracy ranking, all of them shapes that share features with common signs and are therefore easily confused.
The ‘preferred’ versus ‘second’ gap. The thirteen characters with two estimates are exactly the ones the note names as depending on context. The differences the report was willing to state are not small. The Part-Word time for 1-2-4 moves from 143.30 preferred to 206.35 second, a 44 per cent increase in reading time when the sign has to be discriminated against a similar compound. 1-4-5 moves from 143.30 to 206.35 on the same axis. The accuracy figures move in the same direction: 1-3-4 at 96.78 preferred drops to 92.22 second on Part-Word, and the Whole-Word accuracy for 1-3-4 drops from a low 83.72 (already an outlier) to 77.65. These paired numbers are the whole point of the note the report puts at the head of the table: whether Part-Word or Whole-Word usage runs into competing signs is not always a fact about the alphabet, and the report is careful to state both estimates rather than choose.
The residual in the printed ‘Average’. The report closes Appendix B with a four-number row called Average: 153.01, 123.89, 92.56, 96.08. verify.mjs computes the same four means from the recovered values under three subset rules; the closest of them, one row per character preferring the preferred estimate over second and third, gives 154.83, 124.45, 92.00, 95.42. The Whole-Word time and both accuracies land within 0.7 of the printed values, which is inside the rounding tolerance of two-decimal inputs. The Part-Word time is 1.82 units above, which is not.
The residual is not zero, and it is not the result of a transcription error we can find: three re-checks of the highest-value rows against the page image (the four 213.xx entries and the 178.97 pair) match the recovered values, and the same discrepancy holds under every subset rule we tried. What the report used as its denominator, or which subset it averaged, is not reconstructable from the printed values alone. The residual is stated on this page rather than papered over, because the whole point of showing the check is that the check has to be shown when it does not quite close, too.
What this changes
The rankings themselves, published in the next appendix of the same report, still stand: they are ordinal, they were computed from these values, and even the residual in the mean is small enough that the ordering does not move. What changes is what a reader can now do with the report. The numbers behind every ranking are typed into a JSON file and served, at research/tactile-codes/appendix-b/data/table.json in the repository this page ships out of; anyone with an argument about legibility can now sort the table by any column, ask which signs the committee found slowest or least accurate, or check the report’s own arithmetic against its own values. None of that was possible while the archive’s text search returned SHATVA AOVUNOOV.
The method, if you want to run it yourself
The pipeline is short enough to run in a browser tab in five minutes on any machine with curl, imagemagick and tesseract. curl fetches the archive’s derived JPEG for leaf 47; convert -rotate 90 gives it the correct orientation; tesseract --psm 6 tsv writes a TSV of every OCR’d word with its bounding box, which is what decode.mjs uses as an initial pass; verify.mjs runs the three invariants and the four-column cross-check. The hand-transcription step, walking the page image row by row and typing values into JSON, took roughly forty-five minutes and is committed in data/table.json. Every value in this stratum is the same JSON; the browser reads it out of a <script type="application/json"> block at page load, so the table above is not a reprint of a different source, it is the same file the verifier reads.
That kind of provenance is what the third-row layer was already after when it recovered Wait’s 1908 alphabet chart off the pixels three days ago (the OCR text layer for that page had also thrown the information away, and for the same underlying reason: it read what it could name, and could not name the shape of a raised-dot sign in a row and column grid). This page picks the thread up on the same volume, one appendix over, and the same rule applies: when the text layer disagrees with the pixels, the pixels are the source and the text is the derivative, and the derivative is the one to check.
In plain words
The 1915 Uniform Type Committee report printed its per-character legibility values on one landscape sheet inside a portrait book, and every scan puts the page upside-down. The archive's OCR read the headers as 'SHATVA AOVUNOOV' (that is 'ACCURACY VALUES' rotated a hundred and eighty degrees), so the table has been unsearchable ever since. Rotate, re-OCR, and hand-check: 334 legibility values, recovered.