A clock made of words
716 A.D., Give or Take Two Thousand Years
In 1953, Robert Lees counted shared basic words and dated the lexical split between English and German. Reproduce his arithmetic, then change every defensible choice and watch the date move.
First, the published point
Lees printed F = .585, from 124 cognates among 212 compared words, and a retention constant k = .805. His equation was t = ln(F) / 2ln(k).
The paper prints
1.236 millennia
English and German begin diverging about 716 A.D.
Printed 90% band: 470 to 962 A.D.
This browser recomputes
calculating
calculating, residual calculating
Using 124/212 before rounding F gives calculating. Using the paper's printed ±.246 depth gives calculating.
Now turn one date into 1,350
Five word lists, three expert cognacy codings, five retention constants, three formulas, three denominator rules, and two synonym rules are fully crossed. Every control changes one coordinate of the selected cell. The curve still runs all cells.
Your selected specification
Loading the public data
Fetching the 84 KB shipped data file.
calculating. Lees's 716 A.D. is at the calculating, and the calculating.
Blank cells are unsupported no-loan analyses in sources with no loan annotation. They are gaps, never zeroes and never averages.
One ancestor. One elapsed time. Five clocks.
Bergsland and Vogt attacked the idea of a universal lexical clock by comparing descendants of Old Norse. Their paper remains unread here, so none of their reported figures is repeated. This table runs the objection afresh on IE-CoR data. Each daughter is compared with Old Icelandic across the same list and the same 1,020 years.
The Old Icelandic test
| daughter | shared | F | fitted k |
|---|---|---|---|
| Loading |
Strict classified items, any synonym, Swadesh 1950 list. Every value is recomputed from the shipped IE-CoR classifications.
Calculating the seeded permutation check.
Feed the ruler back into the clock
The attested panel fits a separate single-lineage retention rate for each of 14 ancestor-to-descendant pairs. Their mean becomes the “refit” option above. A good average does not make every pair obey it.
Calculating attested-depth errors.
Vacuous corner: this is in-sample calibration, not independent validation. The same 14 pairs estimate the refitted constant and then receive it back. The error spread remains descriptive.
The depths are historical ranges collapsed to Lees-style midpoints: Old English 1.0 millennia, Old High German 1.1, Old Icelandic 1.02, Latin 2.15, and Ancient Greek 2.2. This page does not date Proto-Indo-European.
Which choice moved the answer?
Calculating variance decomposition.
This exact decomposition uses the complete 900-cell common-support grid, where both strict and inclusive rules exist for every source. The 150 defined IE-CoR no-loan cells and 300 unsupported no-loan gaps remain in the full curve above.
What Lees could not have chosen
His sensitivity check changed 716 A.D. to 689 A.D. by using language-specific constants. But IELex arrived in 2021, IE-CoR in 2024, Leipzig-Jakarta in 2009, ASJP in 2008, and Starostin's correction in 2000. The large curve is not a menu Lees concealed. It is the space that accumulated around his honest single calculation.
The check
Two different claims are kept separate. The page exactly reproduces Lees's arithmetic from his printed F and k. It cannot reproduce his cognate count because his lists were not published. The modern curve derives a new input from three later databases.
- The shipped file is a CC BY 4.0 trim of IE-CoR, IELex, Dyen-Kruskal-Black, and five Concepticon lists. It retains only concepts and languages used here, plus ordered cognate assignments and IE-CoR loan flags. Forms themselves are omitted.
- Exactly 300 cells are undefined because IELex and Dyen provide no loan annotation. The coverage rail leaves them blank. Defined coincidences are counted separately, including IE-CoR's identical 1950/1952 intersections and controls that turn out not to move a result.
- IELex's form field contains only ---. Its cognate classes are usable, so presence is determined from form rows and classification records, never from displayed strings.
- The Doubt field is the literal false throughout all three CLDF conversions. A doubt toggle would be dead, so none is offered.
- The three databases are expert opinions with shared intellectual ancestry, not three independent samples of truth. IE-CoR was built for phylogenetics, not this rejected constant-rate dating method. This analysis is ours and does not imply its authors' endorsement.
- The attested-depth layer is narrowed to 14 comparisons in five ancestor groups whose Lees-style midpoint depths are stated here. Proposed Classical Armenian, Old Church Slavonic, and Vedic extensions were omitted because this build did not independently verify defensible date ranges for them. No depth was guessed.
- The concept-specific model fits one penalised hazard per concept from the 14 IE-CoR attested pairs, adding half a retained and half a replaced observation at the mean depth to avoid infinite rates. It then solves the expected sister-pair retention numerically. This is an explicit implementation choice, not a number printed by Lees.
- The .86 constant is labelled as conventionally attributed to Swadesh 1955 because that paywalled paper was not read. Bergsland and Vogt 1962 was also unread and is cited only for its argument, not for any figure.
Sources and redistribution
Lees 1953, sections 3.31 and 4.3; Bergsland and Vogt 1962, argument only; IE-CoR 1.1; IELex Data and Trees; Dyen Indo-European; Concepticon. Data attribution, exact URLs, byte sizes, hashes, trimming, and method are recorded in the research README. All redistributed data is CC BY 4.0.