A census where there was a sample

Tyndale Exactly

In June 1997, at the Library of Congress, William Tyndale's biographer announced a number: 83 percent of the King James New Testament is Tyndale exactly. The study behind it calls itself, in its own subtitle, an estimation based on sampling, and it sampled eighteen passages. Nobody counted the rest. This page counts the rest: every one of the 180,405 words of the King James New Testament, aligned against Tyndale's New Testament of 1534, with the alignment recomputed in your own browser as you read. Counted strictly, word for word and in his order, the answer is near 72 per cent, and the eleven-point gap turns out to be neither an accident of sampling nor a quarrel about Tudor spelling. It is what the word exactly was doing. And the number that matters more is the one nobody had ever put beside it: the Rheims New Testament of 1582, made from the Latin by exiles with every motive to avoid him, still shares 55.7 per cent of Tyndale's words.

There is a strange thing about the sentence a famous number lives in. Daniell's is a good example, and the Library of Congress bulletin that reported it puts the strangeness plainly in the very next line.

David Daniell announced at the Library on June 4 that “83 percent of the King James New Testament is Tyndale exactly.” Although it was known that Tyndale's English translations of the Bible were the basis of the King James Bible, an exact percentage had not been determined, he said.

Yvonne French, "'Courage and Genius': Tyndale Responsible for Most of Kings James Translation, Says Biographer", Library of Congress Information Bulletin, July 1997. loc.gov, read 2026-09-07.

An exact percentage had not been determined, and now one had. The study it rests on appeared the following year: Jon Nielson and Royal Skousen, How Much of the King James Bible Is William Tyndale's? An Estimation Based on Sampling, in Reformation volume 3, pages 49 to 74. It is behind a paywall, its publisher withholds the abstract, and we have not read it. So this page takes no position at all on what they did. It notices only what anybody can notice: the half of their title after the question mark is the half that goes missing, and the figures that travel in its place do not agree with each other.

What is in circulation

Ninety per cent, and one third, in one sentence. Eighty-three per cent, exactly. These are not compatible, and none of them says what a word is, what counts as the same word across four centuries of spelling, or what the denominator is. That is not a scandal. It is what happens to any number that circulates without its definition attached. The only cure is to state a definition, count the whole thing, and show the working, so let us do that.

The measure, stated before it was run

For each verse the King James carries, take Tyndale's stream of words and the King James's stream of words, and find the largest set of King James words that can be matched one to one against Tyndale's without any two matches crossing over each other. That is the longest common subsequence, and it is the honest reading of word for word, in his order. Retention is that total over every King James word in the New Testament, including the words of verses Tyndale's text does not carry, so a verse he never reached counts against him rather than quietly disappearing.

Then there is the hard part, which is the whole of the difficulty. Tyndale wrote In the beginnynge was the worde. The King James prints In the beginning was the Word. Whether that is one word twice or two different words is a judgement, and the entire answer turns on it. So the judgement is made four separate ways, and all four are printed. Switch between them below and watch the number move.

Any verse in the New Testament, aligned in front of you

Red is a King James word that stands in Tyndale. Plain text is a word that does not. Nothing here was precomputed: your browser has the two texts and engine.js, and it is running the alignment now, at the rung you choose.

The words below are shown as the census compares them: lower case, punctuation gone, hyphens split. That is not a tidied quotation, it is the actual object being counted, and showing anything prettier would be showing you something other than the measurement.

The instrument

When are two spellings one word?

  

Recompute the whole thing yourself

One chapter is a demonstration. The claim is about the New Testament, so the button below fetches all twenty-seven books, about two megabytes of text, and runs the same code over all 180,405 words. It takes a few seconds. What it prints is what your machine computed, next to what this repository computed, and whether they agree.

The census, in your browser

Not run yet.

The four rungs, and what each of them could be inventing

A normaliser can be wrong in two directions. It can fail to join two spellings of one word, which loses matches and understates the answer. Or it can join two genuinely different words, which invents matches and overstates it. The second is the dangerous one, and it can be measured exactly, with nobody's judgement involved: take the King James, whose spelling is already modern, so that one spelling is one word, and count how many distinct words each rung collapses onto a shared form. Every collapse is a licence to match wrongly, and the weight of the collapsed groups is a hard ceiling on how much of the answer could be that error.

RungAssumesRetentionKing James word types it fusesShare of its words

Two of the rungs are worth setting against each other, because they share no assumptions. L2 is a theory of Tudor spelling written by hand: four rules, each named, each justified, applied blind. Learned is a count of coincidences: align on the words that already match, and where exactly one unmatched word sits on each side between two matched anchors, record the pair; accept it only if it recurs, only if each form is the other's most frequent partner, only if they are close enough in letters, and only if the two forms are not both native to both texts. That last bar is what keeps a real Tudor merger like then and than out of a table of spellings.

They agree to within about a point, and where they disagree the disagreement has a name. The learner's distance bar rejects sayde against said, two edits over five letters, which the hand rules join. That one pair is 997 occurrences and most of the gap between them.

A sample of the map the texts taught, longest first

Tyndale, 1534King JamesTimes attested

What the number means, which is a separate question from what it is

Seventy-two per cent sounds like near total dependence. Whether it is depends entirely on something nobody had measured: what a translation that is not Tyndale's scores on the same instrument. A retention figure with no baseline is a temperature with no scale.

So here are readings from the same instrument on five different pairs of texts. Two are calibration, and the instrument is only trustworthy because of them: run it on two editions of the King James, which are the same translation, and it should return nearly everything; pair each King James verse with a different Tyndale verse from the same book, holding genre and subject and the underlying Greek as fixed as they can be while breaking the descent between the particular sentences, and it should return very little.

The third is the one that had never been taken. The Rheims New Testament was made from the Latin Vulgate by Catholic exiles in the 1580s, from a church that had burned Tyndale, and it is as close to an independent English rendering of the same book as the century produced. The text used here is the 1582 printing itself, transcribed from the book by the Text Creation Partnership, and not the Douay-Rheims that circulates as public domain, which is Challoner's revision of the 1750s and is known to have drifted toward the King James. Both are on the chart, and the difference between them is worth about three points, which is roughly the size of Challoner's drift.

The same measure, on four different pairs of texts

And the sharper cut. Much of any retention figure is the, and, of. Two English New Testaments of the same Greek will share those whoever wrote them, so a figure that includes them is measuring the English language as much as it is measuring Tyndale. Take out the hundred commonest words of the King James New Testament, which is a rule rather than a choice, and about 36 per cent of the text is left. On that remainder the King James keeps 57.0 per cent of Tyndale, and the 1582 Rheims, made from the Latin by people who had no wish to keep anything of his, still keeps 40.5 per cent.

Why is this number lower? Two experiments, not an opinion

A count that lands eleven points below the figure in circulation owes the reader an account of the gap. There are only two places it can come from. Either the eighteen passages were unrepresentative, or the two are counting different things. Both can be tested, and neither test requires reading the paywalled paper.

First: the same passages, the same instrument

There is one study of this question whose method and passages are published in full. Ronald Mansbridge, writing in the Tyndale Society Journal, states his rule in a single sentence, names his nine chapters, and prints his counts for each one. Six of the nine are in the New Testament, so they can be run through this census unchanged, and any difference is then the counting rule alone, on identical text, with nothing else free to vary.

PassageHis wordsOursHis figureWord for word, in orderOrder ignoredAny near spelling

Second: could eighteen passages have said eighty-three?

Take the finished census, draw eighteen passages at random, form exactly the same ratio, and do it twenty thousand times. Sampling variation is then not an argument, it is a histogram.

Twenty thousand draws of eighteen passages

Four centuries, one man's words

The same measure, run against every public-domain English New Testament in the corpus. Two of them are marked because they are not in the line: the Rheims New Testament was made from the Latin Vulgate by Catholic exiles who had every reason not to borrow from Tyndale, and the Bible in Basic English was written inside a thousand-word vocabulary and could not have borrowed from him even where it wanted to.

TranslationDateLineTyndale retainedContent words only

Where the rest of it came from

Every King James word that is not Tyndale's came from somewhere. Ask the same question of the Geneva Bible of 1599, the version the King James translators had on the desk and were instructed to displace, and the New Testament divides in three.

All 180,405 words of the King James New Testament

Two things the count found that nobody was looking for

The verse numbers on Tyndale's text are not his

He died in 1536. English verse numbering arrives with Whittingham in 1557 and the Geneva Bible in 1560, so every verse number on a Tyndale text was put there by a later editor, and where that editor's division disagrees with the King James's, a verse-level alignment matches the wrong sentences and scores near zero on material that is in fact retained. Removing the verse boundaries and aligning each chapter as one stream is immune to that, so the gap between the two locates every slip without anyone reading a word. Across 260 chapters it found 2.

ChapterAligned verse by verseAligned as one chapter

John 18 is the clear case: from about the fifth verse this digitisation's numbering runs one behind the King James's and rejoins near the end of the chapter. Nothing is missing from the text. Only the numbers slipped, and only a measurement that never trusted them would have noticed.

The King James is not where Tyndale stops being read

The last row of the survival table is a modern revision descended from the American Standard Version, published in 2000. It still holds 52.2 per cent of the words of a man executed in 1536, and 37.3 per cent of them once the hundred commonest words are removed.

Is this really Tyndale?

The whole census rests on one file, and its publisher's copyright page says only The Tyndale New Testament (1534), Public Domain. It names no editor, no base edition and no digitiser. That is not enough to build a number on, so it was checked against two things that publisher had no hand in.

Against an unrelated digitisation

Against the photographed book

The check

What this does not settle