Language · a measurement

A Green Great Dragon

Mark Forsyth wrote that adjectives in English "absolutely have to be in this order", a photograph of the page went round the world in 2016, and it is still taught as a law. So here is the whole of the English Wikipedia asked the question: 4,340,343 noun phrases with two or more adjectives, every pair counted both ways. The preference is real and it is strong. It is also not a law, and it provably cannot be a list: somewhere in the busiest twenty adjectives sit majority preferences that no left-to-right sequence can satisfy at once.

The claim

In 2013 Mark Forsyth published The Elements of Eloquence. Chapter eight, "Hyperbaton", contains this:

John Ronald Reuel Tolkien wrote his first story aged seven. It was about a 'green great dragon'. He showed it to his mother who told him that you absolutely couldn't have a green great dragon, and that it had to be a great green one instead. Tolkien was so disheartened that he never wrote another story for years.

The reason for Tolkien's mistake, since you ask, is that adjectives in English absolutely have to be in this order: opinion-size-age-shape-colour-origin-material-purpose Noun. So you can have a lovely little old rectangular green French silver whittling knife. But if you mess with that word order in the slightest you'll sound like a maniac. It's an odd thing that every English speaker uses that list, but almost none of us could write it out. And as size comes before colour, green great dragons can't exist.

Mark Forsyth, The Elements of Eloquence, Icon Books 2013, chapter 8 ("Hyperbaton"). Transcribed from the Icon Books electronic edition, ISBN 978-1-84831-717-8. The US edition (Berkley 2014, The Elements of Eloquence: Secrets of the Perfect Turn of Phrase) carries the same words with American quotation marks and keeps the British spelling colour. No page number is given here because no paginated copy could be consulted.

On 3 September 2016 Matthew Anderson, then editor of BBC Culture, photographed that page and posted it under ten words of his own, "Things native English speakers know, but don't know we know:", with the list itself visible only in the photograph. Worth being exact, because almost every retelling gets it wrong: the famous sentence was never typed into the tweet. It was a picture of a book. Within a week it had been shared tens of thousands of times and covered by the BBC, NPR, the Washington Post and Quartz, and it has been an English-teaching handout ever since.

It is a lovely claim, and it has the great virtue of being checkable. Word order leaves a trace in every sentence anyone ever wrote down. So this page counts.

Ask the corpus

Every noun phrase in the November 2023 English Wikipedia that begins with a determiner and stacks two or more adjectives in front of a lowercase noun: 6,407,814 articles, 2,852,642,361 words, 4,340,343 such phrases. Every pair of adjectives inside them counted in both of its orders. Type two adjectives and see what English did.

The pair oracle

loading the counts...

The table the oracle reads holds the busiest 62,927 pairs. A pair below that cut is not absent from English, only from this file.

How absolute is "absolutely"?

Take every pair of adjectives that appears at least 30 times in the corpus: 16,924 pairs, 2,094,889 phrases between them. If the order were fixed by a law, each of those pairs would appear in one order and never the other.

8,771 of the 16,924 pairs (51.8%) appear in both orders. Across all 2,094,889 phrases, 95,941 of them (4.58%) sit in the minority order for their own pair. And only 3,688 pairs (21.8%) are lopsided enough that the 95% Clopper-Pearson interval on the majority share clears 0.95, which is the weakest version of "absolutely" worth the word.

Every judged pair, by how lopsided it is

Bars are counts of pairs; the label is the percentage of that pair's phrases that take its commoner order. A law lives in the rightmost bar alone.

Part of that 95% bar is thin data rather than free variation, so here is the same question asked only of pairs the corpus has seen at least 300 times, where thin data cannot be the answer: 1,113 pairs, 937,610 phrases. 704 of them (63.3%) sit above 99 to 1, which is the preference being real. 784 of them (70.4%) nonetheless occur in both orders, which is the preference not being a law: a pair can run 99 to 1 and still have the 1. Across all of them 3.09% of phrases sit in the minority order.

The shape of that histogram is the whole finding in one picture. The preference is real: the mass sits far to the right of 50/50, and no serious reader of English would expect otherwise. But a rule that "absolutely" fixes order predicts a single spike at 100, and that is not what is there.

Sounding like a maniac

Forsyth's test is that breaking the order makes you sound like a maniac. The corpus has the minority orders in it, written by people who were not trying to be strange. Here are sentences from Wikipedia, verbatim, in the order the rule forbids.

This is not a new observation, only a bigger one. Within two days of the tweet Mark Liberman put seven pairs to COCA and found them mixed, writing that "The bit about green great dragons is correct, I think, and it's also correct that there are often strong preferences for prenominal modifier order. But it's not so easy to characterize the preferences". Simon Horobin at Oxford wrote in The Conversation, of the rule, that "it remains an untested, hypothetical, large, sweeping (sorry) claim", breaking it inside the sentence that says so. Michael Brown ran his own COCA searches a fortnight later and wrote of his counterexamples: "None of these constructions sound maniacal." What follows below is the part nobody has done.

The knife

Before the structure, one small thing. The book's own showpiece is "a lovely little old rectangular green French silver whittling knife", eight adjectives in the prescribed order. Nothing that long is ordinary English. Of the 4,340,343 phrases counted here, 3,959,951 stack exactly two adjectives, 351,978 stack three, 26,380 stack four, 1,920 stack five, 104 stack six, 8 stack seven, and 2 stack eight. Simon Horobin made this point in 2016 without the figures; the figure is that a stack of eight turns up 2 times in 4,340,343 phrases. The long ones that do exist are a pleasure: "amateur athletic union swimming meet held at", "north pacific western north american coastal fauna", "french foreign legion free french military personnel", "outward oblique tubercular pale golden metallic fascia". One of those is the rule misfiring in public: WordNet files "at" as a noun, so it can sit in the head slot and let a whole prepositional phrase look like a stack. That is what the 77.5% precision measured below looks like from the inside.

Here is what the corpus has to say about each of the seven adjacent pairs in that phrase.

the book's ordertimesthe other waytimes

The pairs the 2016 argument was actually about

pairWikipediathe books

An encyclopedia is a poor place to look for "big ugly", which is why the second corpus is here. Liberman's 2016 COCA counts, for comparison, ran big/beautiful 43 to 16, little/beautiful 9 to 345, big/ugly 69 to 4, little/ugly 25 to 116, long/tall 18 to 2, big/huge 54 to 20 and big/enormous 2 to 6.

The list that cannot be written

Forsyth says every English speaker uses the list but almost none of us could write it out. There is a stronger reason than forgetfulness. Ask the corpus which of two adjectives comes first and it answers with a majority. Do that for every pair and you get a tournament. A list, any list, is a claim that the tournament is a line.

It is not. Among the 20 busiest adjectives there are 12 three-way cycles out of 861 judged triples: three adjectives A, B, C where the corpus puts A before B, B before C, and C before A. No sequence of words on a page can satisfy all three at once. Not Forsyth's, not Dixon's, not Cinque's, not one that has yet been written.

A cycle built out of three coin flips would prove nothing, so the count was taken again with every edge required to be a majority significant at p < 0.01 by a two-sided exact binomial test. 6 survive. Over the busiest 100 adjectives, where there is more room for three words to disagree, the figures are 745 and 303.

One of them, with its counts

How small a vocabulary does it take? Rank the adjectives by how often they appear in one of these phrases and take the top few. At 4, 6, 8, 10 words the corpus is still consistent with itself. At 12 it is not: the 12 commonest adjectives in the corpus already contain a cycle of significant majorities.

So how wrong is the best possible list?

Every arrangement of adjectives makes some phrases come out backwards. Count them. The 10 words below are the commonest adjectives in the corpus that carry one of the rule's eight classes, at most two from each class, so arranging them is a test of the rule rather than of Wikipedia's subject matter. The corpus holds 14,485 phrases pairing two of them. Arrange them and the counter says how many of those phrases your arrangement calls backwards. The floor is not zero.

Arrange them

Drag a word, or focus one and press the left and right arrow keys to move it. The two buttons below set the two orders worth comparing.

With 10 words the problem is small enough to solve outright, by dynamic programming over all 1,024 subsets, so the best arrangement is a proven minimum and not the best a search happened to find. It gets 1,030 of the 14,485 phrases backwards. The order the rule itself prescribes, opinion before size before age before shape before colour before origin before material, gets 1,478, which is 448 worse than the best any list can do. The two sequences are not the same, and the difference is printed under the words.

Do it for the 20 busiest adjectives in the corpus and the proven minimum is 4,447 phrases out of 125,948 (3.53%). That is the floor for every theory of adjective order that takes the form of a sequence, on this corpus. Of those 4,447, 4,038 are unavoidable even if you were allowed to choose a different order for every pair separately; the remaining 409 are the price of insisting on a single line.

Proven optima at four vocabulary sizes. Every cost is the least any sequence can achieve.
vocabularyphrasesbest possible errorsunavoidable per pairprice of the line3-cycles

The order the corpus draws by itself

Nothing above needed a theory of what adjectives mean. Push the same arithmetic over the 100 busiest adjectives and a sequence falls out, left to right, chosen only to disagree with as few phrases as possible.

Colour it by the eight classes of the rule and the rule is visibly doing real work. Counting only pairs whose two adjectives sit in different classes, 67,622 phrases out of 74,274 (91.0%) are in the order the mnemonic says. Every one of the 6,652 others is a phrase a native speaker wrote and nobody flinched at.

the rule sayswithagainstshare

And 3 of the rule’s own cells point the wrong way. The worst is colour before origin, which the rule requires and the corpus contradicts: 1,239 phrases put colour first and 2,421 put origin first, so the majority runs 66.1% the other way. The others are shape before colour (597 to 643) and shape before origin (43 to 120). This is not a rounding error at the edge of the rule; it is one of the eight arrows in the middle of it, drawn backwards.

The class is not the unit

That 91.0% counts phrases, and a handful of very common pairs can carry a whole cell. Count word pairs instead, one vote each, and the rule holds for 363 of the 396 cross-class pairs the corpus can judge (91.7%). The two figures agree closely, so the cell-level result is not an artefact of a few very common pairs carrying the rest. What it is not is a class-level fact, and the quickest way to see that is below.

The plainest way to see that a class is not acting as a unit is to watch two members of one class disagree about a member of another.

pairfirstsecondthe corpus puts first

Colour before origin is one cell of the rule. Both rows of each pair above are colour and origin. They do not agree, which is what it looks like when the thing doing the ordering is the words and not their classes.

Which adjective belongs to which class is a judgement, and it is mine. The whole assignment is printed in the check below and in research/adjective-order/analyze.py, so you can disagree with it line by line. Nothing else on this page depends on it.

Is that English, or is that Wikipedia?

Two independent tests. First, the same measurement on a different corpus in a different century: 15,660 public-domain books from Project Gutenberg, 927,956,139 words, 908,282 adjective-stacking phrases, mostly written before 1930. Of the 984 pairs common to both corpora with enough data to judge, 947 (96.2%) prefer the same order in the encyclopedia and in the novels.

The registers are not otherwise alike, which is the point of using them both. The busiest pairs in the encyclopedia are "south african", "roman catholic", "north american", "south korean", "english professional", "average annual". In the novels they are "good old", "poor little", "poor old", "little old", "dear old", "dear little".

Second, a test that removes this page's own machinery. The counts above come from a rule with no parser in it: a determiner, four word lists out of WordNet, and adjacency. So the same question was put to 45,369 sentences of Universal Dependencies English treebanks, where a human being has marked every part of speech and every attachment by hand. Where the trees and the rule both see two adjectives modifying one noun, they agree on which came first 729 times out of 730 (99.86%). The trees are also asked the question directly, with no rule involved at all, and where a pair has enough gold instances to have a direction the trees agree with Wikipedia on 8 of 9. That last number is 9 pairs, not 9,000, because hand-parsed English is small: it is worth what a sample of 9 is worth, and no more.

What Tolkien actually wrote

The dragon is worth returning to, because the story as Forsyth tells it is not the story Tolkien told. Here is the whole of it, from a letter to W. H. Auden dated 7 June 1955:

I first tried to write a story when I was about seven. It was about a dragon. I remember nothing about it except a philological fact. My mother said nothing about the dragon, but pointed out that one could not say 'a green great dragon', but had to say 'a great green dragon'. I wondered why, and still do. The fact that I remember this is possibly significant, as I do not think I ever tried to write a story again for many years, and was taken up with language.

J. R. R. Tolkien to W. H. Auden, 7 June 1955. Text as published by the Tolkien Estate at tolkienestate.com. Numbered Letter 163 in Humphrey Carpenter's edition by several secondary sources; no page number is given here because no copy of the printed edition could be consulted, and the 2023 expanded edition repaginates.

The story was about a dragon, not about a green great one. There is no disheartenment in it; Tolkien says only that the memory is "possibly significant". And the man whose childhood is used to prove the rule ends his account of it by saying he never found out why. That is the honest position, and it is still roughly the field's: a century of proposals, from Dixon's seven semantic types to Scott's twenty ordered slots, and the best combination of cognitive predictors yet assembled gets English two-adjective order right about 71% of the time.

He wondered why, and still did. So does everyone. What this page can add is the shape of the thing to be explained: strong, near-universal, gradient rather than categorical, and not a line.

The check

What was counted. A phrase counts when a determiner (one of 15 closed forms) is followed by an optional degree word, then two or more words WordNet 3.1 files as adjectives, then a lowercase word WordNet files as a noun and not as an adjective, whose own next word is not a noun. Text is cut at every character that is not a letter, apostrophe, hyphen or space, so no phrase spans a comma, a bracket or a full stop, and "big and red" cannot match. The lowercase-head condition is what keeps out "the Second World War"; the next-word condition is what keeps out "the lowest round trip fare", where "round" modifies "trip" and not the head at all. Nothing in that rule can see which order the two adjectives are in, and a filter blind to the ordering cannot bias it.

How wrong the rule is. Measured against 45,369 hand-parsed sentences from 15 Universal Dependencies English treebank files. Of the pairs it extracts, 77.5% are genuinely two pre-head modifiers of the same noun, and 62.6% meet the strictest definition (both tagged ADJ, both attached amod). It finds 21.8% of the gold pairs, which is low on purpose: insisting on a determiner throws away most of the language to keep the part that is unambiguous. The errors that remain are mostly nested compounds and participles, and they are the reason the honest word for what is measured here is "prenominal modifier", not "adjective". The one number that could poison the study is order agreement, and it is 729/730 = 99.86%.

The contaminant worth naming. Articles in this Wikipedia dump end with their category and see-also lists as running text, and the rule occasionally reads one of those as a noun phrase: "the Soviet Union Soviet military personnel" is the clearest case. It is measurable, because such a run repeats a word, which real adjective stacks almost never do. Of the 167,097 phrases stacking three or more adjectives, 631 repeat a word, which is 0.015% of all phrases counted. It is left in rather than filtered out, because a filter chosen after seeing the answer is worth less than a number you can subtract yourself.

Free choices, all of them. The degree-word list, the closed-class list and the functional list (ordinals, quantifiers, "other", "same", "such") are declared in extract.py and excluded from the published vocabulary; the eight-class assignment is declared in analyze.py. The pair threshold is 30 instances. Significance is a two-sided exact binomial test with Benjamini-Hochberg control at q = 0.05; 756 of the 16,924 judged pairs do not clear it, which is to say the corpus cannot tell those pairs apart from free variation.

What is new here and what is not. That the rule is gradient rather than categorical is not new, and this page would be dishonest to imply otherwise. Liberman and Horobin said so within days of the tweet. Wulff measured a multifactorial model at 73.5% cross-validated accuracy on the spoken BNC in 2003; Futrell, Dyer and Scontras reached 66% from single predictors on 41,822 held-out triples in 2020; Dyer and colleagues put a century of cognitive predictors together and got 71% on English in 2023, against 85% for a neural network that is explaining nothing. Closest of all, Westbury in 2021 asked 52 native speakers to order 500 attested pairs and published the distribution: only 7% of the pairs were judged unanimously, 9.8% were indeterminate, and he concluded that "These data suggest that the implication sometimes found in the literature that all English prenominal adjectives are strictly ordered is not true.". What is measured here that we could not find measured anywhere is the tournament as a tournament: the cycle count, and the least number of phrases that any single sequence must call wrong. Shaw and Hatzivassiloglou noted in 1999 that both orders can survive transitive closure and arbitrated between them rather than counting; Malouf gave a worked cyclic example from the BNC in 2000; Truswell cited that intransitivity in 2009 as an argument against the cartographic hierarchies without quantifying it; Leung, Emerson and Cotterell forced a total order in 2020 and found it cost their model little accuracy. This page puts a number on that cost in phrases rather than in accuracy. If it has been measured before, we would like to be told, at the door.

Reproduce it. Everything is in research/adjective-order/: the corpora are the pinned wikimedia/wikipedia 20231101.en parquet shards and a Project Gutenberg English mirror, both fetched by fetch.sh; the lexicons come from WordNet 3.1; and node verify-a-green-great-dragon.mjs at the repository root re-derives every number on this page from the committed artifacts, 147 checks, including an independent reimplementation of the extraction rule, of the Clopper-Pearson bounds and of the exact optimum, plus four mutations that must turn it red.