Folklore · a tale-type matrix, recounted

All the Better to Count You With

Jamshid Tehrani coded 58 tellings of Little Red Riding Hood, The Wolf and the Kids and the East Asian Tiger Grandmother on 72 plot traits, and his matrix is public. This page runs it in your browser. 6 of 8 printed network statistics come back at the paper's precision; neither retention index does. Then the paper's own arguments on its own matrix. No trait shared by ATU 333 and ATU 123 is missing from the 15 East Asian tales, it says: as written that is false, since 6 coded states are, but none of those 6 is reconstructed at the two types' common ancestor on any of the 1,000 shortest trees found by this page's bounded search, which does not establish the claim across all 5,740 MPTs. The 6 African and Antiguan tales retain the tested branch with a 5-step search-found gap after this page's 7-trait deletion.

In 2013 the anthropologist Jamshid Tehrani coded 58 tellings of three related stories on 72 plot traits and built family trees from them: Little Red Riding Hood and its relatives (tale type ATU 333 in the international index), The Wolf and the Kids (ATU 123), and the Tiger Grandmother tales of East Asia. His coded matrix is published with the paper. This page runs it, in your browser, and asks it the questions the paper asked.

Every number below is computed here from his supplementary files, except where a number is quoted from the paper, which is marked and cited in place. The page keeps his codes and his wording. Where it adds a reading of its own, it says so and lets you change it.

Tehrani's sentence that ATU 333 and ATU 123 share no trait the East Asian tales lack is false as written on his own matrix. Bounded search: none of 6 is reconstructed at their common ancestor in 1,000 shortest trees found.

I

The matrix

Each row is a tale, each column a trait from Tehrani's File S1, each cell a coded state. A dash is a gap: a trait that could not occur because of something earlier in the story (no voice test, so no operation to clear the voice). His note says gaps are "treated as phylogenetically uninformative by the analyses", and this page treats them the same way unless you say otherwise. The matrix holds 818 gaps and 21 cells marked ?.

state 0 state 1 state 2 higher states in other colours ? (not known) · gap a state File S1 never defines row bar: ATU 333, ATU 123, African, East Asian (the paper's groups)

Choose a tale and a trait using the labelled controls.

File S2 as published, rows grouped by the paper's three groups. The coloured bar at the left of each row is the group the paper's trees put the tale in. Table S1's own "ATU type" column says something different for six of them: it labels 17 tales as ATU 333, 20 as ATU 123 and 21 with a question mark, and those are the 15 East Asian tales and the six African and Antiguan ones. The trees put those six in ATU 123, which makes that group 26. The page keeps the label and the tree group apart throughout.

Two things in the record a reader should know before trusting any tree. The rows for Grimm (1812) and RH1 (Wratislaw's Lusatian "Little Red Hood", 1889) are identical, so no analysis can separate them. And File S1 does not define every state the matrix uses: 7 observed states have no label, among them state 2 of trait 45, "Victim rescued", which 23 tales carry, every one in the paper's ATU 123 group. The page prints each as "state 2, not defined in File S1" and does not guess.

II

Switch off a trait

A parsimony tree is the branching that needs the fewest changes of state to explain the matrix. The shortest this page's search finds needs 330 changes, and it is not one tree: the paper's search "returned 5740 equally most parsimonious trees", and this page's search stops collecting at 1,000 distinct ones. So the page never shows you "the tree". For each group the paper names, it reports two things: whether the group is in every shortest tree the search found, and the gap, which is how many more changes the shortest tree the search can find without that group needs. A gap of 0 means an equally short tree breaks the group up. Switch off a named trait to start a progressive search: early length and group presence arrive first, while the complete gap ledger can take about half a minute.

Traits switched off: none

The control keeps at least one informative trait on. With no informative characters, network distances and parsimony results are undefined, so the page refuses that empty setting.

Showing the full matrix, as computed by this page's engine at its declared seed and shipped with the page.

Group the paper namestalesin the trees foundgap, steps
ATU 333171,0002
ATU 123261,0002
East Asian151,0002
African tales with Antigua61,0001
Catterinella51,0003
The Story of Grandmother41,0001
Little Red Riding Hood71,0001
The Aesopic fables21,0001

Shortest length found 330 · trees collected 1,000 · average delta score 0.2970

On the full matrix the gap is shown as a range over three seeds of the constrained search where they differ. A gap is what this search found, so it is an upper bound: a better search could find a shorter tree without the group.

Strict consensus of the shortest trees found: only branchings present in every one of them, each labelled with its search-found gap. It is drawn rooted on the East Asian group because that is how the paper roots its Table 1, not because the data give a root.

On the full matrix every one of the paper's eight named groups is in all 1,000 trees found, and not one is held by a wide margin: the three main groups by 2, 2 and 2 steps. The paper says ATU 123 "was present in all of the MPTs returned by the original analysis"; it is present in all the trees this search found too, which is agreement about a found set, not about all of them. Switch off trait 5 and the Little Red Riding Hood group, which the paper gave 20% bootstrap support, stops being in every shortest tree: it is in 269 of 300 found, with a gap of 0. The red cap is coded present in 8 tales, and two of the seven in that group (RH3 and Iran) do not have it. Trait 2, "Type of animal", is the one trait that cannot change any tree: switching it off shortens every tree by the same 3 steps and changes no gap, a check you can run.

III

The anchor: numbers that need no tree

Before any argument, the instrument has to reproduce something the paper printed. The paper's NeighbourNet statistics are the cleanest target, because they need no search. For every pair of tales, the page counts the share of traits on which they differ, over the traits both have scored (the uncorrected p-distance). For every four tales, it compares the three ways of pairing them: the delta score of Holland and colleagues (2002) says how far the quartet is from a tree, and the Q-residual does the same on distances rescaled to average 1. The paper computed both in SplitsTree 4.13 and averaged them. There are 424,270 quartets among 58 tales, and 123,410 among the 43 left when the East Asian tales are removed.

A gap against a coded state, in the distance
Quantityprintedcomputed hereat the printed precision
Average delta score, all tales0.30.2970agrees
Average Q-residual, all tales0.030.02833agrees
Average delta score, East Asian removed0.280.2820agrees
Average Q-residual, East Asian removed0.0240.02408agrees
Mean taxon delta, the East Asian tales0.310.3120agrees
Mean taxon Q-residual, the East Asian tales0.040.03543agrees
Mean taxon delta, the other tales0.280.2918does not agree
Mean taxon Q-residual, the other tales0.020.02585does not agree

6 of 8 printed figures reproduce at the precision printed.

Estimates of the overall tree-likeness/boxiness of the network yielded an average delta score of 0.3 and Q-residual score of 0.03.Tehrani 2013, Results
NeighbourNet graph of the data with East Asian tales removed. The average delta score on the Network was 0.28 and the average Q-residual score was 0.024.Tehrani 2013, Figure S3 caption
The average delta score of the East Asian tales is 0.31 compared to an average of 0.28 for the other taxa, while their average Q-residual is 0.04 compared to 0.02.Tehrani 2013, Discussion

Six agree, one of them to three decimal places. That is a real reproduction: SplitsTree's outputs, recovered from nothing but the matrix. The last pair does not. Averaged over the 43 tales that are not East Asian, in the same network, the taxon means are 0.2918 and 0.02585, which print as 0.29 and 0.03. The overall average forces this, because the mean over all 58 is a weighted mean of the 15 and the 43. One reading does reproduce the printed pair, and it is only a reading: the averages of the separate network with the East Asian tales removed, 0.2820 and 0.02408, print as 0.28 and 0.02. If that is what happened, the Discussion compares the East Asian tales in one network with the other tales in another. The page cannot see SplitsTree's session, so it shows both and decides neither. Count a gap against a coded state as a difference and the delta score falls to 0.2714, and only 4 of 8 figures reproduce: the paper's figures fit the skip rule.

IV

What the tree search can say, and what it cannot

The paper's parsimony analysis used PAUP 4 with tree-bisection-reconnection and 1,000 replications. This page's engine does the same kind of search, from scratch, in JavaScript: random addition, tree-bisection-reconnection, and a parsimony ratchet, which briefly reweights a random quarter of the traits to shake a search off a plateau. All traits are unordered and equally weighted, and gaps and ? count as missing. The search is heuristic. A shortest length it reports is the shortest it found.

A gap (-) in the parsimony analysis
AnalysisRI printedshortest foundRI computedminimum mmaximum g
All tales0.723300.7257138838
East Asian tales removed0.762110.7547106534

With gaps as missing data. Retention index RI = (g minus length) / (g minus m).

The cladistic analysis returned 5740 equally most parsimonious trees (MPTs). The fit between the data and the trees was measured with the Retention Index, which was calculated as 0.72.Tehrani 2013, Results
The MPTs returned from the data had Retention Indices of 0.76.Tehrani 2013, Figure S1 caption (the analysis with East Asian tales removed)

Neither retention index comes back. On all the tales the shortest tree found gives 0.7257, which reaches the printed 0.72 only by cutting digits off rather than rounding. With the East Asian tales removed it gives 0.7547 against a printed 0.76, which no rounding reaches. The printed pair would need a longer tree for the first analysis and a shorter one for the second. Removing the 10 traits that only East Asian tales carry, as the paper did, changes neither number, because without those tales the traits are constant. Counting a gap as a state gives 0.7421 and 0.7842, further away still. The page found no convention that reproduces both, and says so rather than choosing one.

Searchseedsshortest found, every seedreplicates reaching it
All tales833043 of 64
East Asian removed821132 of 64
All tales, gap as a state439816 of 32
East Asian removed, gap as a state424816 of 32

Not recomputed, and cited instead: the count of 5740 trees, which depends on PAUP collapsing branches of zero length and on swapping every tied tree to exhaustion; every bootstrap percentage, from PAUP; and every Bayesian posterior probability, from MrBayes 3.2. They belong to programs this page does not run.

V

Three sentences from the Discussion

The sophisticated objection to all of the above is short. Pull any trait out of any matrix and some branch moves; with thousands of tied trees, that tells you about parsimony, not about tales. Fair. So the rest of the page leaves branch-watching behind, takes three claims the paper makes about folklore in its own words, and puts each to the matrix under readings you can change.

A. "Not a single characteristic"

Second, if ATU 333 and ATU 123 are more closely related to each other than they are to the East Asian tales, they would be expected to share derived characters (i.e. novel story traits) that would have evolved after they diverged from the East Asian tradition. However, there is not a single characteristic shared by these two tale types that does not also occur in the East Asian group.Tehrani 2013, Discussion (the second argument against the East Asian tales as a "missing link")

What counts as a characteristic "shared by these two tale types"? The sentence can be read three ways, and the matrix answers each one differently.

Reading of "shared"
Which tales are which type
State 0 (often "absent")
An ancestor with several equally good states
Trait (File S1)statetales with it: ATU 333 / ATU 123 / East Asianthe only state at the common ancestorcounts under this reading
3 Victim is[1] single17 / 9 / 00 of 1,000yes
4 Sex of the victim[2] male2 / 5 / 00 of 1,000yes
9 The relationship of the villain to the victim[3] friend1 / 1 / 00 of 1,000yes
44 The villain falls asleep after the feast[1] present4 / 3 / 00 of 1,000yes
62 Rescued from the villain’s stomach[1] cut out of the monster's belly6 / 11 / 00 of 1,000yes
67 The monster's belly filled with stones[1] present3 / 1 / 00 of 1,000yes

Literal reading, tree groups: 6 coded states are carried by at least one ATU 333 tale and one ATU 123 tale and by none of the 15 East Asian tales. As written, the sentence does not hold.

Read literally, the sentence is false on its own matrix. 6 states are shared by the two types and missing from every East Asian tale: a single victim, a male victim, a villain who is a friend, a villain who falls asleep after the feast, the victim cut out of the belly, and the belly filled with stones. Using Table S1's labels instead of the tree groups still leaves 6. Read as "shared by most tales of each type", it holds: 0 states qualify.

But the paper's own sentence tells you which reading its argument needs. It is about derived characters, "novel story traits" that evolved after the two types split from the East Asian tradition. A state that tales of both types carry today but that their last common ancestor did not have is not that. So the page does what the paper did for its Table 1: it roots each shortest tree found with the East Asian group as the sister of the rest, and asks, trait by trait, which states a most parsimonious reconstruction can put at the last common ancestor of ATU 333 and ATU 123. On all 1,000 trees found by this page's bounded heuristic search, none of the 6 is there. That is a result for this found set, not a reproduction of the paper's claim over all 5740 MPTs. Count state 0 as a trait and one state survives every reading of ancestry: trait 25 at state 0, "The villain’s disguise: [0] absent", the only state at the ancestor on 1,000 trees. The absence of a disguise is hardly a novel story trait, which is why state 0 is off by default and a choice.

B. "Just a few traits"

As mentioned previously, one of the key problems with existing folklore taxonomy is that it defines international types in reference to European type specimens on the basis of just a few traits. In this case, African and East Asian tales are grouped with Little Red Riding Hood because they feature human protagonists, and with The Wolf and the Kids because the villain attacks the victims in their own home, rather than their grandmother's. The phylogenetic approach used here, on the other hand, defines types in reference to the tales' inferred common ancestors rather than any existing variants, and uses all the traits they exhibit as potential evidence for their relationships. This approach yielded clear evidence that the African tales are more closely related to The Wolf and the Kids than they are to Little Red Riding Hood.Tehrani 2013, Discussion

If the trees really use all the traits, the African placement should not hang on the few that carry the index's two criteria. The test: delete those traits, search again, and ask whether the five African tales and the Antiguan one still sit among The Wolf and the Kids tales, meaning that some branch holds all six with at least one ATU 123 tale and no ATU 333 or East Asian tale. The gap is how many more changes the shortest tree found without any such branch needs. Which File S1 traits carry the index's criteria is this page's reading, not Tehrani's, so it is a choice.

The index's traits

With all 72 traits, the African placement holds by 5 steps. With the index's seven traits deleted it holds by 5. Of 40 sets of 7 traits drawn at random from the other informative traits, 36 leave a gap no larger than that, and 0 dissolve the placement.

Gaps for the random deletions (bars) against the index's traits deleted (red line). Each random draw is one search at its own seed.

The placement does not rest on the index's traits. Deleting trait 1 alone leaves a gap of 5; deleting any single one of the other 70 informative traits leaves 4 to 7. Deleting all seven index traits leaves 5, inside the spread of random seven-trait deletions (1 to 7, median 4). That supports the paper's claim, and it is a deletion test the 2013 paper did not report. It also turns up something the paper did not say: with the index's seven traits gone, the full 26-tale ATU 123 group is no longer in every shortest tree (gap 0), and 8 of the 40 random seven-trait deletions also leave that group with a gap of 0. The African tales stay with The Wolf and the Kids while the edges of the group loosen around them.

C. Taking the East Asian tales out

Bootstrap support for the clade separating ATU 333 from ATU 123 increased from 62% to 83%, while the Bayesian posterior probability rose from 87% to 98%.Tehrani 2013, Discussion

The paper removed the 15 East Asian tales to test whether they blend the other two types, and support for the split between ATU 333 and ATU 123 rose. Bootstrap percentages are not recomputed here, but two things are. The delta score falls from 0.2970 to 0.2820. And the gap that holds ATU 333 apart is 2 steps with them and 2 without. Is removing those tales special, or does removing any 15 do it? Of 30 random removals (keeping at least two tales of each type), 0 lower the delta score as far, and 17 leave a gap at least as large.

Average delta score after each random removal (dots) against the East Asian removal (red line).
Gap holding ATU 333 apart after each random removal (bars) against the East Asian removal (red line).

VI

Table 1, recounted

Table 1 is printed as an image, so the page transcribed it (twice, by eye and by optical character recognition, which agree cell for cell) and recounts its "Distribution (n tales)" columns from the matrix. Its "Parsimony (% of MPTs)" column says each trait was absent from the common ancestor of ATU 333 and ATU 123 in every most parsimonious tree.

Counting a tale as having the trait
Trait as printedprinted: East Asian / ATU 123 / ATU 333counted hereagrees
Voice operation (27)2 / 10 / 02 / 10 / 0yes
Hand test (30)8 / 10 / 08 / 10 / 0yes
Dialogue with the villain (32)7 / 0 / 107 / 0 / 10yes
Rescue by passer-by (45)2 / 0 / 71 / 0 / 7no
Excuse to escape (47)9 / 0 / 39 / 0 / 3yes

4 of 5 rows reproduce counting state 1 with the paper's tree groups.

Trait 45 is the exception: File S2 gives 1 East Asian tale rescued by a passer-by (TG11) against the printed 2, and no ? in that column explains the difference. Counting any state above 0 instead reproduces only 2 of 5, and using Table S1's labels only 3 of 5, so the printed table was counted on the tree groups at state 1. The parsimony column reproduces on the trees found here: each of the five traits is absent from the ancestor on every tree, 5 of 5.

The check

Checking the figures in this page against the engine…

A live search on the full matrix starts in your browser when the page loads, and its result is compared with results.json here.

Provenance

Every free choice, and what it moves

A planted fault

Give TG11 state 1 at trait 62 in a copy of the matrix (an East Asian tale cut out of the belly) and run the same engine. The literal list must lose trait 62 and the delta score must move.

What this page will not do

Uncertainties

Sources

Others who have worked on this matrix

David Morrison re-rooted the paper's parsimony consensus and NeighbourNet at their midpoints and re-ran the Bayesian analysis with a relaxed clock, on his blog The Genealogical World of Phylogenetic Networks ("The phylogenetics of Little Red Riding Hood", 4 December 2013). Those three rootings agree in putting the East Asian tales as sister to the rest; a fourth, from a UPGMA tree, does not. Joe Roe reviewed the paper without re-analysing it (2013). Tehrani, Nguyen and Roos returned to the tale's origins with a different sample of 24 versions in Digital Scholarship in the Humanities 31(3), 2016, and Tehrani and d'Huy used the case in Maths Meets Myths (Springer, 2017), whose full text this page could not reach.

We searched PLOS, PubMed Central, the phylogenetic networks blog and general web results on 2026-09-14 and did not find a test of the paper's "not a single characteristic" sentence against its own matrix, or a deletion test of its African placement against random deletions.