Folklore · a tale-type matrix, recounted
All the Better to Count You With
Jamshid Tehrani coded 58 tellings of Little Red Riding Hood, The Wolf and the Kids and the East Asian Tiger Grandmother on 72 plot traits, and his matrix is public. This page runs it in your browser. 6 of 8 printed network statistics come back at the paper's precision; neither retention index does. Then the paper's own arguments on its own matrix. No trait shared by ATU 333 and ATU 123 is missing from the 15 East Asian tales, it says: as written that is false, since 6 coded states are, but none of those 6 is reconstructed at the two types' common ancestor on any of the 1,000 shortest trees found by this page's bounded search, which does not establish the claim across all 5,740 MPTs. The 6 African and Antiguan tales retain the tested branch with a 5-step search-found gap after this page's 7-trait deletion.
In 2013 the anthropologist Jamshid Tehrani coded 58 tellings of three related stories on 72 plot traits and built family trees from them: Little Red Riding Hood and its relatives (tale type ATU 333 in the international index), The Wolf and the Kids (ATU 123), and the Tiger Grandmother tales of East Asia. His coded matrix is published with the paper. This page runs it, in your browser, and asks it the questions the paper asked.
Every number below is computed here from his supplementary files, except where a number is quoted from the paper, which is marked and cited in place. The page keeps his codes and his wording. Where it adds a reading of its own, it says so and lets you change it.
Tehrani's sentence that ATU 333 and ATU 123 share no trait the East Asian tales lack is false as written on his own matrix. Bounded search: none of 6 is reconstructed at their common ancestor in 1,000 shortest trees found.
I
The matrix
Each row is a tale, each column a trait from Tehrani's File S1, each cell a coded state. A dash is a gap: a trait that could not occur because of something earlier in the story (no voice test, so no operation to clear the voice). His note says gaps are "treated as phylogenetically uninformative by the analyses", and this page treats them the same way unless you say otherwise. The matrix holds 818 gaps and 21 cells marked ?.
Choose a tale and a trait using the labelled controls.
Two things in the record a reader should know before trusting any tree. The rows for Grimm (1812) and RH1 (Wratislaw's Lusatian "Little Red Hood", 1889) are identical, so no analysis can separate them. And File S1 does not define every state the matrix uses: 7 observed states have no label, among them state 2 of trait 45, "Victim rescued", which 23 tales carry, every one in the paper's ATU 123 group. The page prints each as "state 2, not defined in File S1" and does not guess.
II
Switch off a trait
A parsimony tree is the branching that needs the fewest changes of state to explain the matrix. The shortest this page's search finds needs 330 changes, and it is not one tree: the paper's search "returned 5740 equally most parsimonious trees", and this page's search stops collecting at 1,000 distinct ones. So the page never shows you "the tree". For each group the paper names, it reports two things: whether the group is in every shortest tree the search found, and the gap, which is how many more changes the shortest tree the search can find without that group needs. A gap of 0 means an equally short tree breaks the group up. Switch off a named trait to start a progressive search: early length and group presence arrive first, while the complete gap ledger can take about half a minute.
Traits switched off: none
The control keeps at least one informative trait on. With no informative characters, network distances and parsimony results are undefined, so the page refuses that empty setting.
Showing the full matrix, as computed by this page's engine at its declared seed and shipped with the page.
| Group the paper names | tales | in the trees found | gap, steps |
|---|---|---|---|
| ATU 333 | 17 | 1,000 | 2 |
| ATU 123 | 26 | 1,000 | 2 |
| East Asian | 15 | 1,000 | 2 |
| African tales with Antigua | 6 | 1,000 | 1 |
| Catterinella | 5 | 1,000 | 3 |
| The Story of Grandmother | 4 | 1,000 | 1 |
| Little Red Riding Hood | 7 | 1,000 | 1 |
| The Aesopic fables | 2 | 1,000 | 1 |
Shortest length found 330 · trees collected 1,000 · average delta score 0.2970
On the full matrix the gap is shown as a range over three seeds of the constrained search where they differ. A gap is what this search found, so it is an upper bound: a better search could find a shorter tree without the group.
On the full matrix every one of the paper's eight named groups is in all 1,000 trees found, and not one is held by a wide margin: the three main groups by 2, 2 and 2 steps. The paper says ATU 123 "was present in all of the MPTs returned by the original analysis"; it is present in all the trees this search found too, which is agreement about a found set, not about all of them. Switch off trait 5 and the Little Red Riding Hood group, which the paper gave 20% bootstrap support, stops being in every shortest tree: it is in 269 of 300 found, with a gap of 0. The red cap is coded present in 8 tales, and two of the seven in that group (RH3 and Iran) do not have it. Trait 2, "Type of animal", is the one trait that cannot change any tree: switching it off shortens every tree by the same 3 steps and changes no gap, a check you can run.
III
The anchor: numbers that need no tree
Before any argument, the instrument has to reproduce something the paper printed. The paper's NeighbourNet statistics are the cleanest target, because they need no search. For every pair of tales, the page counts the share of traits on which they differ, over the traits both have scored (the uncorrected p-distance). For every four tales, it compares the three ways of pairing them: the delta score of Holland and colleagues (2002) says how far the quartet is from a tree, and the Q-residual does the same on distances rescaled to average 1. The paper computed both in SplitsTree 4.13 and averaged them. There are 424,270 quartets among 58 tales, and 123,410 among the 43 left when the East Asian tales are removed.
| Quantity | printed | computed here | at the printed precision |
|---|---|---|---|
| Average delta score, all tales | 0.3 | 0.2970 | agrees |
| Average Q-residual, all tales | 0.03 | 0.02833 | agrees |
| Average delta score, East Asian removed | 0.28 | 0.2820 | agrees |
| Average Q-residual, East Asian removed | 0.024 | 0.02408 | agrees |
| Mean taxon delta, the East Asian tales | 0.31 | 0.3120 | agrees |
| Mean taxon Q-residual, the East Asian tales | 0.04 | 0.03543 | agrees |
| Mean taxon delta, the other tales | 0.28 | 0.2918 | does not agree |
| Mean taxon Q-residual, the other tales | 0.02 | 0.02585 | does not agree |
6 of 8 printed figures reproduce at the precision printed.
Estimates of the overall tree-likeness/boxiness of the network yielded an average delta score of 0.3 and Q-residual score of 0.03.Tehrani 2013, Results
NeighbourNet graph of the data with East Asian tales removed. The average delta score on the Network was 0.28 and the average Q-residual score was 0.024.Tehrani 2013, Figure S3 caption
The average delta score of the East Asian tales is 0.31 compared to an average of 0.28 for the other taxa, while their average Q-residual is 0.04 compared to 0.02.Tehrani 2013, Discussion
Six agree, one of them to three decimal places. That is a real reproduction: SplitsTree's outputs, recovered from nothing but the matrix. The last pair does not. Averaged over the 43 tales that are not East Asian, in the same network, the taxon means are 0.2918 and 0.02585, which print as 0.29 and 0.03. The overall average forces this, because the mean over all 58 is a weighted mean of the 15 and the 43. One reading does reproduce the printed pair, and it is only a reading: the averages of the separate network with the East Asian tales removed, 0.2820 and 0.02408, print as 0.28 and 0.02. If that is what happened, the Discussion compares the East Asian tales in one network with the other tales in another. The page cannot see SplitsTree's session, so it shows both and decides neither. Count a gap against a coded state as a difference and the delta score falls to 0.2714, and only 4 of 8 figures reproduce: the paper's figures fit the skip rule.
IV
What the tree search can say, and what it cannot
The paper's parsimony analysis used PAUP 4 with tree-bisection-reconnection and 1,000 replications. This page's engine does the same kind of search, from scratch, in JavaScript: random addition, tree-bisection-reconnection, and a parsimony ratchet, which briefly reweights a random quarter of the traits to shake a search off a plateau. All traits are unordered and equally weighted, and gaps and ? count as missing. The search is heuristic. A shortest length it reports is the shortest it found.
| Analysis | RI printed | shortest found | RI computed | minimum m | maximum g |
|---|---|---|---|---|---|
| All tales | 0.72 | 330 | 0.7257 | 138 | 838 |
| East Asian tales removed | 0.76 | 211 | 0.7547 | 106 | 534 |
With gaps as missing data. Retention index RI = (g minus length) / (g minus m).
The cladistic analysis returned 5740 equally most parsimonious trees (MPTs). The fit between the data and the trees was measured with the Retention Index, which was calculated as 0.72.Tehrani 2013, Results
The MPTs returned from the data had Retention Indices of 0.76.Tehrani 2013, Figure S1 caption (the analysis with East Asian tales removed)
Neither retention index comes back. On all the tales the shortest tree found gives 0.7257, which reaches the printed 0.72 only by cutting digits off rather than rounding. With the East Asian tales removed it gives 0.7547 against a printed 0.76, which no rounding reaches. The printed pair would need a longer tree for the first analysis and a shorter one for the second. Removing the 10 traits that only East Asian tales carry, as the paper did, changes neither number, because without those tales the traits are constant. Counting a gap as a state gives 0.7421 and 0.7842, further away still. The page found no convention that reproduces both, and says so rather than choosing one.
| Search | seeds | shortest found, every seed | replicates reaching it |
|---|---|---|---|
| All tales | 8 | 330 | 43 of 64 |
| East Asian removed | 8 | 211 | 32 of 64 |
| All tales, gap as a state | 4 | 398 | 16 of 32 |
| East Asian removed, gap as a state | 4 | 248 | 16 of 32 |
Not recomputed, and cited instead: the count of 5740 trees, which depends on PAUP collapsing branches of zero length and on swapping every tied tree to exhaustion; every bootstrap percentage, from PAUP; and every Bayesian posterior probability, from MrBayes 3.2. They belong to programs this page does not run.
V
Three sentences from the Discussion
The sophisticated objection to all of the above is short. Pull any trait out of any matrix and some branch moves; with thousands of tied trees, that tells you about parsimony, not about tales. Fair. So the rest of the page leaves branch-watching behind, takes three claims the paper makes about folklore in its own words, and puts each to the matrix under readings you can change.
A. "Not a single characteristic"
Second, if ATU 333 and ATU 123 are more closely related to each other than they are to the East Asian tales, they would be expected to share derived characters (i.e. novel story traits) that would have evolved after they diverged from the East Asian tradition. However, there is not a single characteristic shared by these two tale types that does not also occur in the East Asian group.Tehrani 2013, Discussion (the second argument against the East Asian tales as a "missing link")
What counts as a characteristic "shared by these two tale types"? The sentence can be read three ways, and the matrix answers each one differently.
| Trait (File S1) | state | tales with it: ATU 333 / ATU 123 / East Asian | the only state at the common ancestor | counts under this reading |
|---|---|---|---|---|
| 3 Victim is | [1] single | 17 / 9 / 0 | 0 of 1,000 | yes |
| 4 Sex of the victim | [2] male | 2 / 5 / 0 | 0 of 1,000 | yes |
| 9 The relationship of the villain to the victim | [3] friend | 1 / 1 / 0 | 0 of 1,000 | yes |
| 44 The villain falls asleep after the feast | [1] present | 4 / 3 / 0 | 0 of 1,000 | yes |
| 62 Rescued from the villain’s stomach | [1] cut out of the monster's belly | 6 / 11 / 0 | 0 of 1,000 | yes |
| 67 The monster's belly filled with stones | [1] present | 3 / 1 / 0 | 0 of 1,000 | yes |
Read literally, the sentence is false on its own matrix. 6 states are shared by the two types and missing from every East Asian tale: a single victim, a male victim, a villain who is a friend, a villain who falls asleep after the feast, the victim cut out of the belly, and the belly filled with stones. Using Table S1's labels instead of the tree groups still leaves 6. Read as "shared by most tales of each type", it holds: 0 states qualify.
But the paper's own sentence tells you which reading its argument needs. It is about derived characters, "novel story traits" that evolved after the two types split from the East Asian tradition. A state that tales of both types carry today but that their last common ancestor did not have is not that. So the page does what the paper did for its Table 1: it roots each shortest tree found with the East Asian group as the sister of the rest, and asks, trait by trait, which states a most parsimonious reconstruction can put at the last common ancestor of ATU 333 and ATU 123. On all 1,000 trees found by this page's bounded heuristic search, none of the 6 is there. That is a result for this found set, not a reproduction of the paper's claim over all 5740 MPTs. Count state 0 as a trait and one state survives every reading of ancestry: trait 25 at state 0, "The villain’s disguise: [0] absent", the only state at the ancestor on 1,000 trees. The absence of a disguise is hardly a novel story trait, which is why state 0 is off by default and a choice.
B. "Just a few traits"
As mentioned previously, one of the key problems with existing folklore taxonomy is that it defines international types in reference to European type specimens on the basis of just a few traits. In this case, African and East Asian tales are grouped with Little Red Riding Hood because they feature human protagonists, and with The Wolf and the Kids because the villain attacks the victims in their own home, rather than their grandmother's. The phylogenetic approach used here, on the other hand, defines types in reference to the tales' inferred common ancestors rather than any existing variants, and uses all the traits they exhibit as potential evidence for their relationships. This approach yielded clear evidence that the African tales are more closely related to The Wolf and the Kids than they are to Little Red Riding Hood.Tehrani 2013, Discussion
If the trees really use all the traits, the African placement should not hang on the few that carry the index's two criteria. The test: delete those traits, search again, and ask whether the five African tales and the Antiguan one still sit among The Wolf and the Kids tales, meaning that some branch holds all six with at least one ATU 123 tale and no ATU 333 or East Asian tale. The gap is how many more changes the shortest tree found without any such branch needs. Which File S1 traits carry the index's criteria is this page's reading, not Tehrani's, so it is a choice.
With all 72 traits, the African placement holds by 5 steps. With the index's seven traits deleted it holds by 5. Of 40 sets of 7 traits drawn at random from the other informative traits, 36 leave a gap no larger than that, and 0 dissolve the placement.
The placement does not rest on the index's traits. Deleting trait 1 alone leaves a gap of 5; deleting any single one of the other 70 informative traits leaves 4 to 7. Deleting all seven index traits leaves 5, inside the spread of random seven-trait deletions (1 to 7, median 4). That supports the paper's claim, and it is a deletion test the 2013 paper did not report. It also turns up something the paper did not say: with the index's seven traits gone, the full 26-tale ATU 123 group is no longer in every shortest tree (gap 0), and 8 of the 40 random seven-trait deletions also leave that group with a gap of 0. The African tales stay with The Wolf and the Kids while the edges of the group loosen around them.
C. Taking the East Asian tales out
Bootstrap support for the clade separating ATU 333 from ATU 123 increased from 62% to 83%, while the Bayesian posterior probability rose from 87% to 98%.Tehrani 2013, Discussion
The paper removed the 15 East Asian tales to test whether they blend the other two types, and support for the split between ATU 333 and ATU 123 rose. Bootstrap percentages are not recomputed here, but two things are. The delta score falls from 0.2970 to 0.2820. And the gap that holds ATU 333 apart is 2 steps with them and 2 without. Is removing those tales special, or does removing any 15 do it? Of 30 random removals (keeping at least two tales of each type), 0 lower the delta score as far, and 17 leave a gap at least as large.
VI
Table 1, recounted
Table 1 is printed as an image, so the page transcribed it (twice, by eye and by optical character recognition, which agree cell for cell) and recounts its "Distribution (n tales)" columns from the matrix. Its "Parsimony (% of MPTs)" column says each trait was absent from the common ancestor of ATU 333 and ATU 123 in every most parsimonious tree.
| Trait as printed | printed: East Asian / ATU 123 / ATU 333 | counted here | agrees |
|---|---|---|---|
| Voice operation (27) | 2 / 10 / 0 | 2 / 10 / 0 | yes |
| Hand test (30) | 8 / 10 / 0 | 8 / 10 / 0 | yes |
| Dialogue with the villain (32) | 7 / 0 / 10 | 7 / 0 / 10 | yes |
| Rescue by passer-by (45) | 2 / 0 / 7 | 1 / 0 / 7 | no |
| Excuse to escape (47) | 9 / 0 / 3 | 9 / 0 / 3 | yes |
4 of 5 rows reproduce counting state 1 with the paper's tree groups.
Trait 45 is the exception: File S2 gives 1 East Asian tale rescued by a passer-by (TG11) against the printed 2, and no ? in that column explains the difference. Counting any state above 0 instead reproduces only 2 of 5, and using Table S1's labels only 3 of 5, so the printed table was counted on the tree groups at state 1. The parsimony column reproduces on the trees found here: each of the five traits is absent from the ancestor on every tree, 5 of 5.
The check
Checking the figures in this page against the engine…
A live search on the full matrix starts in your browser when the page loads, and its result is compared with results.json here.
Provenance
- SHA-256 of every data file is recomputed in your browser and compared with record.json.
Every free choice, and what it moves
- Gap in the distance: skip the pair (default, reproduces the anchor) or count a difference. Moves the delta score.
- Gap in parsimony: missing data (default, File S1) or an extra state. Moves lengths and retention indices.
- Which tales are which type: tree groups (default) or Table S1 labels. Moves the shared-state counts.
- Reading of "shared": literal (default), majority or ancestral. Moves which states count.
- State 0 as a trait: no (default) or yes. Adds trait 25 at state 0.
- Ambiguous ancestral sets: strict (default) or any. Inert on this record: on the trees found, none of the tallied states is ever one of several equally good states at the common ancestor, so counting ambiguous sets changes no count.
- Table 1 counting rule: state 1 (default) or any non-zero state. Moves the recount.
- The index's traits for B: trait 1, or trait 1 with traits 10, 13, 16, 22, 24, 25 (default). Moves the search length and the ATU 123 gap.
- Search effort and seeds are declared in engine.js (PLAN) and shipped in results.json; the seed spread is shown in section IV.
A planted fault
Give TG11 state 1 at trait 62 in a copy of the matrix (an East Asian tale cut out of the belly) and run the same engine. The literal list must lose trait 62 and the delta score must move.
What this page will not do
Uncertainties
- Every tree result is heuristic: the shortest length found, the trees collected before a cap, and gaps that are upper bounds. Section IV shows the seed spread.
- The retention indices do not reproduce, and the page found no convention that reproduces both.
- The "other taxa" network figures do not reproduce; the reading that does is a guess about the software session.
- Table 1 row 45 disagrees with the matrix by one tale.
- File S1 leaves 7 observed states undefined, prints [1] twice for trait 64, and breaks off mid-label for trait 41.
- The mapping of the index's criteria to File S1 traits is this page's reading.
- The coding itself is Tehrani's, from English translations; this page cannot check a single cell against a tale.
Sources
- Jamshid J. Tehrani (2013), "The Phylogeny of Little Red Riding Hood", PLoS ONE 8(11): e78871, doi:10.1371/journal.pone.0078871. Article, with Table S1, File S1 and File S2, retrieved 2026-09-14. Licensed CC BY 3.0: "provided the original author and source are credited". Changes: text extracted from the DOCX supplements; Table 1 transcribed; no cell altered.
- B. R. Holland, K. T. Huber, A. Dress and V. Moulton (2002), "δ Plots: A Tool for Analyzing Phylogenetic Distance Data", Molecular Biology and Evolution 19: 2051 to 2059, and R. D. Gray, D. Bryant and S. J. Greenhill (2010), "On the shape and fabric of human history", Philosophical Transactions of the Royal Society B 365: 3923 to 3933: the two sources the paper cites for the delta score and the Q-residual.
- Kevin C. Nixon (1999), "The Parsimony Ratchet, a new method for rapid parsimony analysis", Cladistics 15.
Others who have worked on this matrix
David Morrison re-rooted the paper's parsimony consensus and NeighbourNet at their midpoints and re-ran the Bayesian analysis with a relaxed clock, on his blog The Genealogical World of Phylogenetic Networks ("The phylogenetics of Little Red Riding Hood", 4 December 2013). Those three rootings agree in putting the East Asian tales as sister to the rest; a fourth, from a UPGMA tree, does not. Joe Roe reviewed the paper without re-analysing it (2013). Tehrani, Nguyen and Roos returned to the tale's origins with a different sample of 24 versions in Digital Scholarship in the Humanities 31(3), 2016, and Tehrani and d'Huy used the case in Maths Meets Myths (Springer, 2017), whose full text this page could not reach.
We searched PLOS, PubMed Central, the phylogenetic networks blog and general web results on 2026-09-14 and did not find a test of the paper's "not a single characteristic" sentence against its own matrix, or a deletion test of its African placement against random deletions.