A census where there was a sample
Tyndale Exactly
In June 1997, at the Library of Congress, William Tyndale's biographer announced a number: 83 percent of the King James New Testament is Tyndale exactly. The study behind it calls itself, in its own subtitle, an estimation based on sampling, and it sampled eighteen passages. Nobody counted the rest. This page counts the rest: every one of the 180,405 words of the King James New Testament, aligned against Tyndale's New Testament of 1534, with the alignment recomputed in your own browser as you read. Counted strictly, word for word and in his order, the answer is near 72 per cent, and the eleven-point gap turns out to be neither an accident of sampling nor a quarrel about Tudor spelling. It is what the word exactly was doing. And the number that matters more is the one nobody had ever put beside it: the Rheims New Testament of 1582, made from the Latin by exiles with every motive to avoid him, still shares 55.7 per cent of Tyndale's words.
There is a strange thing about the sentence a famous number lives in. Daniell's is a good example, and the Library of Congress bulletin that reported it puts the strangeness plainly in the very next line.
David Daniell announced at the Library on June 4 that “83 percent of the King James New Testament is Tyndale exactly.” Although it was known that Tyndale's English translations of the Bible were the basis of the King James Bible, an exact percentage had not been determined, he said.
Yvonne French, "'Courage and Genius': Tyndale Responsible for Most of Kings James Translation, Says Biographer", Library of Congress Information Bulletin, July 1997. loc.gov, read 2026-09-07.
An exact percentage had not been determined, and now one had. The study it rests on appeared the following year: Jon Nielson and Royal Skousen, How Much of the King James Bible Is William Tyndale's? An Estimation Based on Sampling, in Reformation volume 3, pages 49 to 74. It is behind a paywall, its publisher withholds the abstract, and we have not read it. So this page takes no position at all on what they did. It notices only what anybody can notice: the half of their title after the question mark is the half that goes missing, and the figures that travel in its place do not agree with each other.
What is in circulation
Ninety per cent, and one third, in one sentence. Eighty-three per cent, exactly. These are not compatible, and none of them says what a word is, what counts as the same word across four centuries of spelling, or what the denominator is. That is not a scandal. It is what happens to any number that circulates without its definition attached. The only cure is to state a definition, count the whole thing, and show the working, so let us do that.
The measure, stated before it was run
For each verse the King James carries, take Tyndale's stream of words and the King James's stream of words, and find the largest set of King James words that can be matched one to one against Tyndale's without any two matches crossing over each other. That is the longest common subsequence, and it is the honest reading of word for word, in his order. Retention is that total over every King James word in the New Testament, including the words of verses Tyndale's text does not carry, so a verse he never reached counts against him rather than quietly disappearing.
Then there is the hard part, which is the whole of the difficulty. Tyndale wrote In the beginnynge was the worde. The King James prints In the beginning was the Word. Whether that is one word twice or two different words is a judgement, and the entire answer turns on it. So the judgement is made four separate ways, and all four are printed. Switch between them below and watch the number move.
Any verse in the New Testament, aligned in front of you
Red is a King James word that stands in Tyndale. Plain text is a word that does not. Nothing here was
precomputed: your browser has the two texts and
engine.js, and it is running the alignment
now, at the rung you choose.
The words below are shown as the census compares them: lower case, punctuation gone, hyphens split. That is not a tidied quotation, it is the actual object being counted, and showing anything prettier would be showing you something other than the measurement.
The instrument
When are two spellings one word?
Recompute the whole thing yourself
One chapter is a demonstration. The claim is about the New Testament, so the button below fetches all twenty-seven books, about two megabytes of text, and runs the same code over all 180,405 words. It takes a few seconds. What it prints is what your machine computed, next to what this repository computed, and whether they agree.
The census, in your browser
Not run yet.
| Rung | What it treats as one word | Your browser | This repository | Agrees |
|---|
The four rungs, and what each of them could be inventing
A normaliser can be wrong in two directions. It can fail to join two spellings of one word, which loses matches and understates the answer. Or it can join two genuinely different words, which invents matches and overstates it. The second is the dangerous one, and it can be measured exactly, with nobody's judgement involved: take the King James, whose spelling is already modern, so that one spelling is one word, and count how many distinct words each rung collapses onto a shared form. Every collapse is a licence to match wrongly, and the weight of the collapsed groups is a hard ceiling on how much of the answer could be that error.
| Rung | Assumes | Retention | King James word types it fuses | Share of its words |
|---|
Two of the rungs are worth setting against each other, because they share no assumptions. L2 is a theory of Tudor spelling written by hand: four rules, each named, each justified, applied blind. Learned is a count of coincidences: align on the words that already match, and where exactly one unmatched word sits on each side between two matched anchors, record the pair; accept it only if it recurs, only if each form is the other's most frequent partner, only if they are close enough in letters, and only if the two forms are not both native to both texts. That last bar is what keeps a real Tudor merger like then and than out of a table of spellings.
They agree to within about a point, and where they disagree the disagreement has a name. The learner's distance bar rejects sayde against said, two edits over five letters, which the hand rules join. That one pair is 997 occurrences and most of the gap between them.
A sample of the map the texts taught, longest first
| Tyndale, 1534 | King James | Times attested |
|---|
What the number means, which is a separate question from what it is
Seventy-two per cent sounds like near total dependence. Whether it is depends entirely on something nobody had measured: what a translation that is not Tyndale's scores on the same instrument. A retention figure with no baseline is a temperature with no scale.
So here are readings from the same instrument on five different pairs of texts. Two are calibration, and the instrument is only trustworthy because of them: run it on two editions of the King James, which are the same translation, and it should return nearly everything; pair each King James verse with a different Tyndale verse from the same book, holding genre and subject and the underlying Greek as fixed as they can be while breaking the descent between the particular sentences, and it should return very little.
The third is the one that had never been taken. The Rheims New Testament was made from the Latin Vulgate by Catholic exiles in the 1580s, from a church that had burned Tyndale, and it is as close to an independent English rendering of the same book as the century produced. The text used here is the 1582 printing itself, transcribed from the book by the Text Creation Partnership, and not the Douay-Rheims that circulates as public domain, which is Challoner's revision of the 1750s and is known to have drifted toward the King James. Both are on the chart, and the difference between them is worth about three points, which is roughly the size of Challoner's drift.
The same measure, on four different pairs of texts
And the sharper cut. Much of any retention figure is the, and, of. Two English New Testaments of the same Greek will share those whoever wrote them, so a figure that includes them is measuring the English language as much as it is measuring Tyndale. Take out the hundred commonest words of the King James New Testament, which is a rule rather than a choice, and about 36 per cent of the text is left. On that remainder the King James keeps 57.0 per cent of Tyndale, and the 1582 Rheims, made from the Latin by people who had no wish to keep anything of his, still keeps 40.5 per cent.
Why is this number lower? Two experiments, not an opinion
A count that lands eleven points below the figure in circulation owes the reader an account of the gap. There are only two places it can come from. Either the eighteen passages were unrepresentative, or the two are counting different things. Both can be tested, and neither test requires reading the paywalled paper.
First: the same passages, the same instrument
There is one study of this question whose method and passages are published in full. Ronald Mansbridge, writing in the Tyndale Society Journal, states his rule in a single sentence, names his nine chapters, and prints his counts for each one. Six of the nine are in the New Testament, so they can be run through this census unchanged, and any difference is then the counting rule alone, on identical text, with nothing else free to vary.
| Passage | His words | Ours | His figure | Word for word, in order | Order ignored | Any near spelling |
|---|
Second: could eighteen passages have said eighty-three?
Take the finished census, draw eighteen passages at random, form exactly the same ratio, and do it twenty thousand times. Sampling variation is then not an argument, it is a histogram.
Twenty thousand draws of eighteen passages
Four centuries, one man's words
The same measure, run against every public-domain English New Testament in the corpus. Two of them are marked because they are not in the line: the Rheims New Testament was made from the Latin Vulgate by Catholic exiles who had every reason not to borrow from Tyndale, and the Bible in Basic English was written inside a thousand-word vocabulary and could not have borrowed from him even where it wanted to.
| Translation | Date | Line | Tyndale retained | Content words only |
|---|
Where the rest of it came from
Every King James word that is not Tyndale's came from somewhere. Ask the same question of the Geneva Bible of 1599, the version the King James translators had on the desk and were instructed to displace, and the New Testament divides in three.
All 180,405 words of the King James New Testament
Two things the count found that nobody was looking for
The verse numbers on Tyndale's text are not his
He died in 1536. English verse numbering arrives with Whittingham in 1557 and the Geneva Bible in 1560, so every verse number on a Tyndale text was put there by a later editor, and where that editor's division disagrees with the King James's, a verse-level alignment matches the wrong sentences and scores near zero on material that is in fact retained. Removing the verse boundaries and aligning each chapter as one stream is immune to that, so the gap between the two locates every slip without anyone reading a word. Across 260 chapters it found 2.
| Chapter | Aligned verse by verse | Aligned as one chapter |
|---|
John 18 is the clear case: from about the fifth verse this digitisation's numbering runs one behind the King James's and rejoins near the end of the chapter. Nothing is missing from the text. Only the numbers slipped, and only a measurement that never trusted them would have noticed.
The King James is not where Tyndale stops being read
The last row of the survival table is a modern revision descended from the American Standard Version, published in 2000. It still holds 52.2 per cent of the words of a man executed in 1536, and 37.3 per cent of them once the hundred commonest words are removed.
Is this really Tyndale?
The whole census rests on one file, and its publisher's copyright page says only The Tyndale New Testament (1534), Public Domain. It names no editor, no base edition and no digitiser. That is not enough to build a number on, so it was checked against two things that publisher had no hand in.
Against an unrelated digitisation
Against the photographed book
The check
- The corpus is twelve public-domain translations, fetched as USFM from eBible.org.
research/tyndale-census/sources.jsonrecords the SHA-256 of every archive, andnode research/tyndale-census/fetch.mjs --checkre-downloads them, refuses any whose digest has moved, and compares the parse against the committed text. - The measure, the four rungs, the learner's acceptance criteria and every control were fixed in
research/tyndale-census/PREREGISTRATION.md, which was committed together with the four modules and beforecensus.mjsexisted. The git history is the evidence. - Each module carries its own controls, including ones designed to fail:
node research/tyndale-census/usfm.mjs --selftest,node research/tyndale-census/normalize.mjs --selftest,node research/tyndale-census/lcs.mjs --selftest,node research/tyndale-census/learn.mjs --selftest. - The census itself is
node research/tyndale-census/census.mjs, about a minute, no network once the corpus is on disk. The provenance checks arenode research/tyndale-census/provenance.mjs --check; the 1582 Rheims is fetched and parsed bynode research/tyndale-census/rheims1582.mjs; the replication of the published study isnode research/tyndale-census/mansbridge.mjs; and what this page is handed is written bynode research/tyndale-census/emit.mjs. - From an empty directory, owning none of this. Every program above is served at
/checks/under its own repository path. The corpus is not: a check is a program and the data it reads is a dataset, sodata/text/*.tsvandsources.jsondo not go there.fetch.mjsrebuilds the corpus from eBible.org's own archives and refuses any whose SHA-256 has moved, and it needs the registry, so this page serves it. Nine files, no repository, and about two minutes:
That was run in a directory holding nothing else, and the figures it printed are the ones above.mkdir aw && cd aw for f in usfm normalize lcs learn fetch rheims1582 census; do curl -sS --create-dirs -o research/tyndale-census/$f.mjs \ https://artwaste.land/checks/research/tyndale-census/$f.mjs done curl -sS --create-dirs -o research/what-lincoln-said/align.mjs \ https://artwaste.land/checks/research/what-lincoln-said/align.mjs curl -sS -o research/tyndale-census/sources.json \ https://artwaste.land/strata/tyndale-exactly/data/sources.json node research/tyndale-census/fetch.mjs # 12 witnesses, each digest-checked node research/tyndale-census/rheims1582.mjs # the 1582 Rheims, from the TCP transcription node research/tyndale-census/census.mjs # prints every figure on this page - Everything this page states is asserted by
node verify-tyndale-exactly.mjs, which reads the built page, recomputes the census independently, and checks that the code your browser just ran returns the same answer as the code in the repository over all 180,405 words.
What this does not settle
- We have not read the paper. It is paywalled and its abstract is withheld. Nothing here is a correction of Nielson and Skousen. The resampling experiment shows only that sampling variation cannot carry a figure computed this way up to eighty-three, which means the two are counting different things, and says nothing about which counting is better.
- The Old Testament is untouched. Tyndale translated the Pentateuch and Jonah and left Joshua to Chronicles in manuscript. None of that is measured here.
- The independence arm is close to clean, not perfectly clean. The 1582 Rheims is a
genuinely separate English rendering, but its translators knew the Protestant Bibles and argued with
them, and both texts descend from a shared Latin and Greek tradition, so some of the
55.7 per cent is common ancestry rather than independence. Four editorial
decisions were needed to read the TEI transcription at all, and each is stated at the top of
research/tyndale-census/rheims1582.mjs; the largest is that the compositor set w as a doubled v, which had to be folded. Philemon has no division of its own in that transcription and was recovered by its running head. - The middle group is not the Geneva's invention. Coverdale 1535, Matthew 1537, the Great Bible 1539 and the Bishops' Bible 1568 all stand between Tyndale and Geneva, and none of them was available as a public-domain machine-readable text. Read that share as entered the line after Tyndale and by 1599, not as anyone's authorship.
- One King James, in modern spelling. Three editions were run and the answer moved by four hundredths of a point, so the choice does not matter at this precision. But no original-spelling 1611 text was found, so the question of what the 1769 revision did to Tyndale's share is open.
- A retention figure is not a measure of debt. Keeping a word is not the same as taking it, two translators of one Greek sentence will often reach for the same English, and this instrument cannot tell those apart. That is exactly why the baseline arms are on the page and not in a footnote.