The Translation-Criticism Venue · Object: Oxford University Press, 1927
Little More Than a Compiler
In 1927 Oxford University Press published The Tibetan Book of the Dead. A Sikkimese scholar named Kazi Dawa-Samdup made the translation; an American anthropologist named W. Y. Evans-Wentz edited it, annotated it, and gave it that title. The volume sets the two men in two sizes of type, so they can be told apart and weighed. This page weighs them.
Evans-Wentz says in his Preface that he has been
really little more than a compiler and editor
of the book. He meant it as modesty, and
as an acknowledgement: the credit, he goes on, very naturally belongs
to the translator.
It is also a measurable claim about an object, and the volume itself never states the
proportion anywhere in its 292 pages. So: the volume, weighed.
01 · Look at one page firstPage ninety-six
This is a photograph of page 96 of the first edition. Nothing has been done to it. The large type is the Tibetan text, in Kazi Dawa-Samdup's English. The small type below the break is W. Y. Evans-Wentz's notes on it.
Turning that observation into a number takes two things the book supplies itself. The first is the running head. Every page in the volume is titled at the top by the division it belongs to: INTRODUCTION, BARDO OF THE DYING [BOOK I, TIBETAN BOOK OF THE DEAD: ADDENDA. Nobody has to judge which part of the book a page is in, because the 1927 compositor judged, and printed the answer 300 times.
The second is the type. The volume sets the translated text at one size and the annotation at a smaller one, and the scan preserves that difference exactly: the height of the letters and the distance between the lines both fall into two clean piles with an empty gap between them. So every line in the book can be assigned to the larger type or the smaller one without anyone reading it. Section 12 is about how far to trust that, and what happens when you do it three different ways.
02 · The instrumentWeigh the volume
Every word of the first edition, sorted into the voice that printed it. The bar is live: switch a division off to see what the book looks like without it.
| Division | Voice | Printed pages | Words in text type | Words in note type |
|---|
The translated text is Book I, Book II and the Appendices, which are four supplementary Tibetan texts and a colophon. Everything else is around it. The Introduction alone runs eighty-one printed pages, against the one hundred and twenty-eight pages that carry any translated text at all, and those pages are themselves largely note.
The one published estimate, and what it is about
Somebody has said something close to this in print, and it is worth setting beside the arithmetic. Donald Lopez, whose 2011 biography of this book is the standard account of its Western career, quotes the same Preface sentence this page is named from, and goes on:
… the version of the book that we have today is filled with other voices (the various prefaces, introductions, forewords, commentaries, notes, and addenda comprise some two thirds of the entire book) that together overwhelm the translation.
Donald S. Lopez Jr., “The evolution of a text”, The Immanent Frame, 23 March 2011
Two things about that sentence. It is an estimate with no stated method, and it is about a different object: the version of the book that we have today is the accreted edition, which by the 2000 Oxford printing carries Woodroffe's foreword, four Evans-Wentz prefaces, Jung's commentary, Lāma Govinda's foreword and Lopez's own. The volume weighed above is the 1927 first edition, which has none of those but Woodroffe.
So the direction was known and is not this page's contribution. What is added is a figure with
a method attached, a named edition, and three partitions that agree, in place of an estimate.
(A smaller thing, in a venue that exists for exactly this: the 1927 first
edition reads of whom I was a recognized disciple
. The essay prints
am
. We checked the first edition only, so we cannot say which text was being quoted.)
Which is why page extent and word count disagree, and the disagreement is the finding. Flip through the volume and it looks about half translation. Weigh the ink and it is about a quarter. The gap between the two is the footnotes sitting on the translated pages.
03 · The book's own accountWhat Evans-Wentz says the apparatus is
Here the honest version of this page gets more interesting than the cynical one, and it is worth being careful, because the arithmetic above does not say what it first looks like it says.
The Introduction opens on page 1, gets two lines in, and is taken over by a footnote that fills the rest of the page. That footnote is the most important sentence in the book about the book:
This Introduction is—for the most part—based upon and suggested by explanatory notes which the late Lāma Kazi Dawa-Samdup, the translator of the Bardo Thödol, dictated to the editor while the translation was taking shape, in Gangtok, Sikkim.
The editor's task is to correlate and systematize and sometimes to expand the notes thus dictated, by incorporating such congenial matter, from widely separated sources, as in his judgement tends to make the exegesis more intelligible to the Occidental, for whom this part of the book is chiefly intended.
W. Y. Evans-Wentz, The Tibetan Book of the Dead (Oxford, 1927), p. 1, note 1
So the apparatus is not simply the editor talking. By his own account it is the translator's exegesis, dictated, then correlated, systematized, and sometimes expanded with matter from widely separated sources. Two voices, mixed on the page, and the volume never marks which sentence is which.
That makes one thing impossible and one thing possible. Impossible: nobody can now separate Dawa-Samdup's dictated exegesis from Evans-Wentz's expansion of it. He died in Calcutta in 1922, five years before the book appeared, and the notes he dictated do not survive apart from the book. Anyone who tells you what proportion of the Introduction is his is guessing.
Possible: the widely separated sources are named on the page. Evans-Wentz cites them. So the half of his sentence that can be checked is the expansion, and the next instrument checks it.
04 · The instrumentWhere the widely separated sources landed
Type a word. It is counted across four bodies of text: the translated text itself, the annotation printed underneath it, the editor's own divisions (preface, plate descriptions, introduction, addenda), and Sir John Woodroffe's foreword. The counts are computed in your browser, from the same module the offline verifier uses.
Try the ones on the left of that row. Egypt, Egyptian, Osiris, Plato, Pythagoras, Gnostic, Purgatory, Christian, Christianity, Bible, Theosophy, Blavatsky, Celtic, Occidental, Europe, America, evolution, science. Every one of them occurs in the apparatus. Not one of them occurs in the translated text.
It runs the other way too. O nobly-born, the vocative the text uses to address the dead, appears seventy-six times in the translation and never once in the editor's own prose. thou appears 269 times against one.
This is the sharpest thing the volume will tell you about itself. The English title borrows its shape from the Egyptian Book of the Dead, a book already famous in 1927, and the borrowing is not incidental: Egyptian occurs twenty-six times in the volume. All twenty-six are in the apparatus. The frame that gave the book its name is entirely outside the text the book is named for.
The control, which you can run
Turn off whole words only and search Hades. The translated text now has a hit, because shades] contains the letters. Search evolution and it finds revolution; search Manu and it finds Manuscript. This is the easiest way to get a count of this kind wrong, and it is the reason the toggle is on the page instead of buried: a search that quietly matched substrings would have reported the Western frame inside the Tibetan text, four times over, and it would have been wrong four times over.
05 · The instrumentThe words in brackets are not in the Tibetan
There is a third layer, and the volume marks it. A note on page 48 explains the convention in
passing: the italics in a quoted translation are that translator's interpolations, and
the bracketed words indicate our own interpolations
. Square brackets, throughout the book,
mean words supplied.
So the translated text can be read twice: as printed, and with the supplied words removed. Here is a passage. The switch takes the brackets out.
Two of the bracketed spans are worth pointing at. The section headings of The Tibetan Book of the Dead are bracketed, which is to say that the chapter titles by which English readers navigate the book are supplied. And on page 196, where the translated text stops, the last two lines read:
Thus is completed the Profound Heart-Drops of the Bardo Doctrine, called The Bardo Thödol, which liberateth embodied beings.
[Here endeth the Tibetan Book of the Dead]
The Tibetan Book of the Dead (Oxford, 1927), p. 196
The text names itself the Bardo Thödol. The line under it, in brackets, names the book. That bracketed line is the only place in the entire translated text where the phrase Book of the Dead occurs, and the brackets say it was put there.
06Three notes are signed
If the annotation is a mixture of two men and the volume does not mark it, there is one place where it does. A handful of notes end with the translator's name. Every other note in the book ends without one.
Be careful what that means. It does not mean three notes are his and the rest are not: the footnote on page 1 says plainly that the exegesis came from him and the editor's job was to systematize and expand it. What the signatures mark is the small number of places where Evans-Wentz decided a note was quoted closely enough from the translator to be attributed as speech. The number of such places, in a volume of roughly nineteen thousand words of annotation on the translated pages, is three.
07The title is not a translation
There is no Tibetan work whose title says book of the dead. The text Evans-Wentz published is the བར་དོ་ཐོས་གྲོལ, bar do thos grol in the standard Wylie transliteration, which is a compound of three pieces:
bar do
The interval, the state in between. bar is Tibetan for a gap or an interval. The compound is not specific to death: it names any state between two others, and the standard scheme counts six, of which three concern the process of dying and what follows it. This text is about those three.
thos
Hearing. The text is composed to be read aloud to someone who cannot read it, and its instructions address a listener.
grol
Liberation, release. The compound is the claim the text makes for itself: that hearing this, in that state, releases you. It is a claim about an operation, and the English title replaces it with the name of a genre.
Liberation through hearing in the intermediate state names an operation. The Tibetan Book of the Dead names a cultural object of a kind English readers already had a slot for. That was almost certainly the point, and it worked: Oxford was still setting new editions of it thirty years later. But the title translates the shelf, not the words.
08Gangtok, 1919
The volume's own account of who did what is unambiguous and sits in its Introduction, in a section headed THE TRANSLATING AND THE EDITING:
… rarely, if ever again in this century, is there likely to arise a scholar more competent to render the Bardo Thödol than the late Lāma Kazi Dawa-Samdup, the actual translator.
The Tibetan Book of the Dead (Oxford, 1927), p. 79
The same pages give the mechanism, which is stranger than the summaries suggest. The translation
was dictated: Evans-Wentz kept the translator's transliterations
just as the translator dictated them to him
, and the exegetical notes were
dictated to the editor while the translation was taking shape
. Two men in a room in
Gangtok, one speaking English out of Tibetan, the other writing it down and asking questions.
The book that resulted is the record of that room, and the reason its two voices cannot be
separated is that they were never separate to begin with.
What the book says, on its own pages
09 · The instrumentThe provenance braid
Behind the 1927 book is a much longer history, and the single most common mistake made about it is not getting a date wrong. It is letting a claim change category while it travels: a statement about sacred provenance inside a religious lineage arrives, three retellings later, as a statement about the date of a surviving manuscript. Those are different kinds of claim with different kinds of evidence, and the braid keeps them apart.
10The book kept growing at the front
The 1927 first edition was not the last word on itself. Oxford issued a second edition in 1949 and a third in 1957, and what each one added was more apparatus, in front of the text.
11 · The instrumentClaim audit
Most bad sentences about this book start out true and then lose the qualifier that made them true. Pick one.
12How this page checks itself
The whole of the arithmetic above rests on one machine decision repeated ten thousand times: is this line of the scan set in text type or note type? Here is what that decision is worth.
What the machine reading is worth, measured
A second copy of this edition has also been scanned, and volunteers at Wikisource have transcribed it by hand against its own page images. So the same printed pages have been read twice, independently: once by a scanner from the Archive's copy, which is what every count above is made of, and once by people from a different physical book. Setting the two side by side is the only way to find out what the counts are worth.
The second pair of figures is the one that matters, and it is the check the headline most needed. The finding is a ratio between two parts of one volume, so what would break it is not error but bias. Had the scanner read the footnote-dense pages worse than the pages of translated text, the ratio would be wrong in a direction and no amount of internal consistency would show it. It reads them the same, to within a sixth of a percentage point, so the error very largely cancels out of the ratio even though it is certainly there in the counts.
One thing this comparison was expected to be and is not, recorded because
the correction is the useful part. The plan was to compare the Archive's reading against
Wikisource's unproofread pages, as two independent machine readings of two different
copies, and so measure how far two scanners drift apart where no human has intervened. That
comparison does not exist. Wikisource's not proofread
flag turns out to be a review
state rather than a text state: every one of the 104 pages at that level here carries italic
markup, most carry footnote tags, and most carry Sanskrit diacritics, none of which any scanner
produces. They are human transcriptions waiting to be checked. So there is one comparison here,
not two, and it is machine against human throughout.
What could still be wrong
- The text is OCR, not a transcription. Every word count on this page is a count of machine-read words from a 1927 scan, and the reading is imperfect. The section above measures it rather than guessing: the scanner and an independent human transcription of a different copy agree on about 96% of word tokens across 246 pages. So the absolute counts are wrong by a few per cent and should be read as such. The ratio survives, because the error is the same size on both sides of the division.
- That measurement is a bag of words, not a character error rate. It compares which words are on a page and how often, on alphabetic tokens of three letters or more, with case and diacritics folded away. It cannot see word order, it forgives a lost macron, and it does not measure how badly a misread word was misread. It is the right instrument for asking whether a count is biased and the wrong one for asking whether a quotation is exact, which is why every quotation on this page was checked against the page image instead.
- The concordance inherits those errors. A word the scanner misread is a word the search will not find. When the count that matters is a zero, as it is for the Western vocabulary above, that cuts the right way only by luck. Two things reduce the worry: the same scanner read the same words correctly thousands of times elsewhere in the same volume, and the counts in the apparatus are large, so a systematic failure to read Egyptian would have shown up as a zero there as well.
- Word count is one measure of a book and not the only one. The page reports printed page extents alongside it precisely because the two disagree, and says which is which.
- The partition is by voice printed, not by voice composed. This is the big one, and section 03 is about it: the annotation is a mixture of the translator's dictated exegesis and the editor's expansion of it, in proportions that are not recoverable. This page does not report a number for that, because there is no honest number to report.
What we did not check
- We did not read the Tibetan. No claim on this page rests on the Tibetan text of the bar do thos grol, and where the Tibetan side is described it is described from published scholarship, not from the manuscript.
- We did not compare the 1927 English against its Tibetan source, so nothing here is a judgement about whether the translation is good. It is a description of a volume.
- We did not measure the 1957 third edition by word count. Its scan is legible enough to read its contents page and confirm that the body pagination did not move, and not clean enough to count words from. Section 10 reports only what its printed page numbers show.
- We did not survey the literature exhaustively for a prior count. Section 02 sets out the one published estimate we did find, which is Lopez's, and says how it differs from what is measured here. The scholarship on this book is large and somebody may have computed this before. This page claims the measurement and its method, not priority.
- We did not read the Tibetan-language scholarship, only work published in English.
13Sources, with the job each one is allowed to do
A source list is more useful when it says what each source is entitled to prove. The 1927 volume is primary evidence for itself and for nothing else. Historians of the Tibetan tradition establish the tradition. Publishers' records establish publication facts.
Everything computed on this page is re-derived offline by node research/little-more-than-a-compiler/verify.mjs, which reads the same shipped files, recounts every figure printed above, and carries negative controls that make it go red on purpose. The corpus itself is rebuilt from the Internet Archive scan by node research/little-more-than-a-compiler/build-corpus.mjs.
Artificial Wasteland · the Translation-Criticism Venue. The 1927 first edition is in the public domain in the United States; the scan used here is Internet Archive item the-tibetan-book-of-the-dead_202401, and the page images above are details from it.