Latin epigraphy · information theory · Roman demography

Cut Only What They Could Not Guess

A Roman tombstone is mostly a form letter. You can measure exactly how much of one, because the stonecutter left the evidence in the stone: he carved in full only the words a reader could not have supplied, and reduced the rest to initials. The same instinct governs the ages. A grown man's years are rounded away to nothing. An infant's are counted in hours.

26,108 Latin epitaphs · Epigraphic Database Heidelberg, CC BY-SA 4.0

One stone, twice

This is HD000001 in the Epigraphic Database Heidelberg: a veined marble tablet, thirty-three centimetres by thirty-four, found near Cumae on the bay of Naples, cut somewhere between 71 and 130 AD. Here is what is actually on it.

as carved

Twenty-one words. Now here is the same stone as an epigrapher reads it, with every abbreviation opened out. The pale letters are the ones the editor supplies. They are not on the marble.

as read

D M is Dis Manibus, to the spirits of the dead. Two letters stand for eleven. P F is Publi filiae, daughter of Publius: two letters for eleven again. The names, though, are cut out in full. Nonia Optata. Gaius Iulius Artemo. Nobody could have guessed those.

That contrast is not an accident of this one tablet. It is the strongest regularity in the whole corpus, and you can watch it hold across twenty-six thousand stones.

What belongs to everyone

Below is a real epitaph, drawn from 229 stones packed into this page. The slider asks a single question of every word on it: how many of the other Latin epitaphs also carry this word? Slide right and the words shared with more and more stones fade out. What is left at the far right is what belongs to this stone alone.

Gold marks a word that the database's own prosopography records as part of a name on this stone. Warm red marks a word that is neither a name nor shared above the threshold. Pale grey letters are editorial: either the expansion of an abbreviation or a restoration inside a break in the stone.

Push the slider to the left and almost everything survives, because almost every word is rare enough. Push it right and the epitaph collapses into a name and a number. That collapse is the ordinary case, not the exception: the hundred commonest words account for of every word carved in the entire corpus.

The commonest words on a Latin epitaph, with the share of stones carrying each, how often it was carved as an abbreviation, and what fraction of its letters reached the stone at all.
wordstonesshareabbreviatedletters carved

The law in the stone

Sort every one of the corpus's word types by how many stones carry it, and ask how much of the word the mason actually cut. The answer climbs steadily across nine bands spanning four orders of magnitude.

carried byword typestimes carvedabbreviatedletters carved

A word that appears on exactly one stone in the corpus is abbreviated of the time, and of its letters reach the marble. A word carried by more than three thousand stones is abbreviated of the time, and only of its letters are cut. The mason spent his chisel on exactly what his reader could not supply.

The climb is unbroken through the first eight bands and then dips slightly in the ninth, and the reason is worth stating rather than hiding, because it is the same effect the next table controls for. The ninth band holds only word types, and its average word is shorter than the eighth band's ( letters against ). One of those eight is et, which appears on stones and is abbreviated of the time, because a two-letter word has nothing to give up. Drop the words under four letters from both bands and the order comes back: .

The obvious objection is length: common words are short, short words are less worth abbreviating, so perhaps this is a fact about length wearing a costume. It is not. Hold word length fixed and the effect is still there, in every single row.

Percentage of occurrences carved as an abbreviation, by word length (down) and by the number of stones sharing the word (across). Cells with fewer than 40 occurrences are left blank.
letters

Zipf argued in 1949 that frequent words are short because speakers economise, and Piantadosi, Tily and Gibson refined that in 2011 by showing a word's average information content in context predicts its length better than raw frequency does. What is on these stones is Zipf's version caught in a different medium. The Latin words did not shorten. The carving did, in proportion to how little news it carried.

The first name is formula too

The clearest case is a Roman's own name, which has two slots that behave in opposite ways. The praenomen, the personal first name, is drawn from a tiny closed set: distinct forms appear across people on these stones, and ten of them cover of all of them. The cognomen is wide open: distinct forms across people, with the ten commonest covering only .

So the word that sounds most like a person, the first name, is the most predictable thing on the stone, and the mason treats it accordingly. Praenomina in the text are abbreviated of the time, an average of letters actually cut. Cognomina are abbreviated of the time, an average of letters cut. Gaius was a C. Artemo was Artemo.

Does the chisel track information, or only length?

That contrast suggests a sharper question. A word's surprisal is how many bits it carries: rare words are surprising, common ones are not. Does surprisal predict how much of a word the mason cut, or is the effect just that rare words are longer and so have more letters available to drop?

Mean letters carved is the wrong thing to measure, because it is capped by the word's length outright: you cannot carve six letters of a five-letter word. The carved fraction has no such ceiling, and on a logit scale it turns out to be very nearly independent of length, which explains of it. Surprisal explains , and with length alongside it the fit reaches while surprisal keeps its whole coefficient. Frequency is estimated on a held-out half of the stones, so the number a word is scored with is not read off the same carvings it is used to predict.

One number cuts against that and belongs here rather than in a footnote. Weight each word type equally instead of each carved word, and the same relation explains only . The strong version is a fact about the words a reader actually meets, which are the common ones. The vocabulary as a list barely shows it. Both are true and they answer different questions, so both are printed.

A caveat on the pedigree, since it would be easy to overclaim: surprisal here is a unigram estimate, and -log2(count/total) is an exact affine transform of log frequency. So this measures Zipf's 1949 law of abbreviation, not the 2011 refinement, which turns on contextual predictability and would need a conditioning model this corpus does not yet have.

The coincidence, tested and dropped

An earlier version of this page noticed that the two name slots carve almost the same number of letters per bit, for the praenomen and for the cognomen, and flagged it honestly as something nobody had tested. Testing it took two goes, and the first one was wrong in a way worth showing.

The lexicon regression above cannot settle it. Its surprisal is marginal, a word's rarity in the corpus as a whole; the name-slot figure is conditional, the entropy of what could have stood in that slot given you are in it. Those are different quantities, and the two published points cannot even be plotted on the regression's axes. Reporting the regression as a verdict on the coincidence would have been a category error, and this page did exactly that for about an hour after it first went up.

The right test compares slots to slots. Below are positions in the epitaph's grammar where a small set of alternatives competes for the same job, each with the entropy of what could have gone there and the letters actually cut.

slotformstokensbitsletters cutper bit

It is not constant. Setting aside the one slot whose entropy is so near zero that any ratio to it explodes, the remaining run from to letters per bit, a spread of , with a coefficient of variation of against a threshold of agreed before the numbers were seen. The two name slots sit at the bottom of that range, next to each other and next to the third name slot, because they are all name slots. Two members of one family resembling each other was never evidence of a law covering the rest.

What survives is the weaker and better claim, and it is the one the whole page rests on: the chisel tracks how much a reader could not have guessed, and not merely how long the word was.

The number nobody knew

April 297: 35. April 308: 37. August 308: 40. Before June 309: 45. June 309: 40. The recorded ages of Aurelius Isidorus, an Egyptian landowner, over twelve years. Duncan-Jones opens his 1977 paper with them. Isidorus grows younger.

Roman epitaphs record age at death, and the ages are not real. Here is the last digit of every age between 23 and 62 in the corpus. There are ten digits. If ages were known, each would take about a tenth of the total.

of recorded adult ages are multiples of five. Duncan-Jones built an index for exactly this in 1977: take the percentage of ages divisible by five separately within each decade from 23 to 62, average the four, subtract 20 and multiply by 1.25, so that perfect accuracy scores 0 and total rounding scores 100. This corpus scores .

He computed it on different data. Italian civilian male citizens came out at 42.8; Italian town-councillors, the most educated group in the evidence, at 15.1; freedmen and slaves at 49.5; males across Africa and Numidia at 51.4. Our higher figure is what his own result predicts, because the Heidelberg database's stated geographic focus is the provinces of the Roman Empire rather than Italy and Rome, and he found rounding worst furthest from the centre. Within our own data the same gradient appears.

provinceages in spanrounding index

He also observed that rounding intensifies with age, which is why he insisted on a decade-by-decade analysis rather than one lump. That reproduces here, cleanly, in a corpus assembled decades after he wrote: . The older you were, the less anybody knew.

What these ages are not. They are not a mortality record. Keith Hopkins settled this in 1966: the distribution of ages at death on Roman tombstones is, in his words, "demographically most improbable", because who received a stone and who had an age carved on it were both culturally patterned. His conclusion was blunt: these ages "must be discarded as useful evidence for estimating life expectation." Nothing on this page estimates a life expectancy, and the median recorded age here () is a fact about commemoration, not about dying.

The number the army kept

Rounded ages could mean two quite different things. Either people genuinely did not know how old they were, or five was simply the polite unit, a stylistic habit like saying "about fifty". Those are hard to tell apart, and a mere pile of rounded numbers cannot do it.

Duncan-Jones found the discriminating test, and buried it in a footnote. Soldiers' epitaphs carry two numbers: age at death, which nobody had written down anywhere, and years of service, which the army recorded. Same men. Same stones. Same masons. Same convention, if convention is what this is.

On stones in this corpus that record years of service, the service figures over the span 3 to 27 are divisible by five of the time, against the 20% you would expect from numbers nobody rounded. The ages on those same stones are divisible by five of the time. Duncan-Jones, working from Levison's nineteenth-century tabulations rather than from this database, reported 23.4% at Rome, 21.7% in Germany and 22.1% in Africa for the same span, and called the deviation "almost negligible".

Our figure sits above his, and we are not going to smooth that over: 28.8% is a real excess, not the near-nothing he found, and it varies by province in our data from 23.3% in Germania superior to 34.1% in Dalmatia. What survives intact is the comparison the control was built to make. The number the bureaucracy held is close to unrounded. The number that lived only in somebody's memory is rounded almost beyond use, on the same stone, about the same man. It was not a style. They did not know.

And then the hours

The same corpus contains the opposite of ignorance. An epitaph could give an age in years, or years and months, or down to days, or down to hours. Ask which lives got the finer units.

age at deathpeoplemonths givendays givenhours given

Part of that gradient is mechanical and we should say so plainly: you cannot describe a six-month-old in whole years, so months are forced. But the forcing runs out early, and the gradient does not. At 20 to 39 the stones still give days of the time; past 40 it falls to . Precision is not a property of the age. It is a property of how closely somebody was watching.

At the far end of that are people in the whole corpus whose stone counts the hours. Their median age is years. of them had not reached five.

counted in hours

Somewhere behind each of those numbers a person sat down and worked out, or simply knew, how many hours it had been. The same culture that could not tell you whether a grown man was fifty or fifty-five counted a child's life to the hour. The stone records precisely what could not be assumed, and precision is what attention leaves behind.

The numbers, recomputed

Everything above is recomputed in your browser from the packed corpus, and independently offline from the raw source. The offline verifier re-parses the original EpiDoc XML with a second implementation written separately from the first, in a different language, and fails if the two disagree on any figure.

not yet run

Where this is soft, and what it is not

Data: Epigraphic Database Heidelberg, Open Data Repository, CC BY-SA 4.0 (edh.ub.uni-heidelberg.de/data), EpiDoc XML dump and prosopography table, as published 7 December 2021.  R. P. Duncan-Jones, "Age-rounding, Illiteracy and Social Differentiation in the Roman Empire", Chiron 7 (1977) 333-354.  K. Hopkins, "On the probable age structure of the Roman population", Population Studies 20:2 (1966) 245-264, doi:10.1080/00324728.1966.10406097.  G. C. Whipple, Vital Statistics: an Introduction to the Science of Demography (2nd edn, 1923) 180-181, the source of the 23-62 span.  G. K. Zipf, Human Behavior and the Principle of Least Effort (1949).  S. T. Piantadosi, H. Tily and E. Gibson, "Word lengths are optimized for efficient communication", PNAS 108:9 (2011) 3526-3529, doi:10.1073/pnas.1012551108.