The Verification Venue · a method that fails harder the more you feed it
More Data, More Certain, Still Wrong
Here is a four-species tree. We know it is the truth, because we are about to grow the DNA on it ourselves. Then we hand the DNA to a reconstruction method and ask for the tree back. In one region of branch-length space the method returns a tree that never existed, and returns it harder the more sequence you give it.
Every phylogeny you have ever seen is an inference. Nobody watched the branchings. What makes the inference trustworthy is supposed to be a guarantee called statistical consistency: as the data grow without bound, the estimate converges on the truth. In 1978 Joe Felsenstein showed that maximum parsimony, then the dominant method, does not have that guarantee. Worse than not having it: there are trees for which parsimony converges on a specific wrong answer, with the support for that wrong answer growing steadily as you add sites.
The bench below is that proof, made operable. The tree on the left is real by construction. Drag its branch lengths. The verdict panel computes, exactly and instantly, what parsimony would say if you handed it infinitely many sites.
Units are expected substitutions per site. Two long branches that are not neighbours, separated by a short internal branch: that is the whole recipe. Presets:
With unlimited data, parsimony converges on
computing
Parsimony's whole decision rests on three numbers: the probability that a site shows two taxa with one base and two with another. Only 36 of the 256 possible site patterns separate the three topologies, and these are they. The bars are exact probabilities, not samples.
Zone boundary, this short/internal
·
long branches past this length flip the verdict
Expected tree length, true topology
·
steps per site, AB|CD
Expected tree length, AC|BD
·
parsimony picks the smaller
The map below is the whole parameter plane, coloured by that same exact calculation: which topology wins the informative-pattern count when the data are unlimited. The pale region is safe. The dark region is what the literature calls the Felsenstein zone. Your current settings are the ring.
Now grow the data
The exact verdict says what happens at infinity. Here is what happens on the way there. For each sequence length we simulate whole alignments down the true tree under Jukes-Cantor, score all three topologies by Fitch parsimony, and record how often the true one wins outright. The blue curve is a safe tree (long branches 0.25). The red curve is the same tree with one number changed: the long branches lengthened to 0.45.
True tree recovered at 100 sites
·
of 120 replicates
At 10 000 sites
·
of 120 replicates
Change
·
a hundredfold more data
Read the red curve again, slowly. It is not noisy, and it is not flat. It goes down. A hundredfold increase in evidence takes the method from occasionally right to reliably wrong, and it would keep going: the wrong topology's informative-pattern probability exceeds the true one's by a fixed margin, so the error is a law of large numbers working perfectly, on the wrong quantity. Bootstrap support for the wrong tree climbs toward 100% for exactly the same reason, resampling the sites cannot escape what the sites say. That last point is not computed on this bench; it is a known result, shown by simulation in Hillis & Bull, Syst. Biol. 42:182-192 (1993), and it follows from the bootstrap proportion tracking the same asymptotic winner. Confidence measures repeatability, not truth.
The mechanism has a name, long-branch attraction. Two lineages that have each changed a great deal will, by chance alone, sometimes land on the same base at the same site. Parsimony reads shared bases as shared ancestry, and cannot tell a coincidence from an inheritance. The two long branches are given many chances to coincide; the short internal branch, which carries the real signal, is given almost none.
This is not only a simulation result. In the 1990s microsporidia (fast-evolving parasites with reduced genomes) were placed at the base of the eukaryote tree by early rRNA analyses, and were later reassigned to the fungi once the long-branch artefact was recognised and denser taxon sampling and better models were applied. The Ecdysozoa hypothesis, that arthropods and nematodes share a moulting clade, was resisted in part because the fast-evolving nematode Caenorhabditis was being pulled toward other long branches. Both cases are now settled; they are here as history, not as open questions.
The answer everyone gives
If you have taken a phylogenetics course you already know the reply, and it is correct: parsimony is naive, use maximum likelihood, which is consistent. So let us do that, on exactly the data that just broke parsimony. Same simulation, same zone, but the three topologies are now scored by fitting Jukes-Cantor branch lengths and comparing log-likelihoods.
| sites | parsimony | likelihood (JC) |
|---|
Same seeds, same alignments, same branch lengths (long 0.45, short 0.05, internal 0.05). Forty replicates per row.
Likelihood wins, decisively, and it wins because more data helps it. That is consistency, demonstrated rather than asserted. The lecture could end here, and usually does.
Where consistency actually lives
It should not end here, because the guarantee has a precondition that gets dropped in the retelling. Likelihood is consistent when the model contains the truth. The theorem is about the pair (method, model), and the word "consistent" attaches to the pair. Take the truth out of the model and the guarantee goes with it.
So here is a process the Jukes-Cantor tree model does not contain. Real sites do not all evolve at the same relative rate down every lineage; the pattern changes over time, a phenomenon called heterotachy. We build the simplest possible version: the true topology stays AB|CD throughout, but a fraction ρ of sites evolve with A and C on long branches, and the remaining sites evolve with B and D on the long branches instead. Every site is generated on the true tree. Nothing is faked. The alignment simply has no single set of branch lengths that describes all of it.
Three methods now compete on that alignment: parsimony, likelihood fitting one homogeneous Jukes-Cantor tree (the misspecified model), and likelihood fitting a two-class mixture of branch lengths with the class weight free (a model that does contain the generating process). Slide ρ and run.
ρ = 0 or 1 is a plain homogeneous tree. ρ = 0.5 is the 50:50 split the published exchange was fought over.
Short branches 0.10, internal branch 0.20, 1000 sites, 60 replicates.
At ρ near 0 or 1 there is no heterotachy, the homogeneous model is correct, and likelihood is far ahead. As ρ approaches 0.5 the misspecification bites, and the blue curve falls below the green one: on this bench, at that setting, parsimony beats likelihood. Parsimony did not improve. Its curve is nearly flat across the whole range: it is not good here, it is merely unbothered, because it never made a claim about the process in the first place. What changed is that likelihood's claim became false.
And the purple curve is the point. Give likelihood a model that contains the heterogeneity and it is back on top wherever the heterogeneity is real, without changing the method at all. Read the endpoints honestly, though: at ρ = 0 and ρ = 1 there is no heterotachy to model, and the mixture pays for its extra parameters, finishing 0.0833 and 0.0333 below plain Jukes-Cantor likelihood on this bench. That is the ordinary cost of fitting a richer model than the data needs, and it is small next to the 0.8833 the plain model loses at ρ = 0.50 by being wrong. Consistency was never a property of "likelihood". It was a property of a model that happened to be right.
The part that is still an argument
That green-beats-blue crossing is not a curiosity we invented. It is the subject of a sharp published exchange, and the exchange is worth laying out because the disagreement is about exactly the quantity you were just sliding.
Evolutionary rates vary among sites and across the phylogenetic tree (heterotachy). A recent analysis suggested that parsimony can be better than standard likelihood at recovering the true tree given heterotachy. Spencer, Susko & Roger, Mol. Biol. Evol. 22:1161 (2005), opening lines
- The claim. Kolaczkowski & Thornton, Nature 431:980-984 (2004), simulated four-taxon alignments concatenated from two different branch-length regimes and reported that maximum parsimony generally outperformed maximum likelihood and Bayesian methods across much of the parameter space they examined. Their abstract puts it as "Maximum parsimony performs substantially better than current parametric methods over a wide range of conditions tested", and that phrasing is the whole fight: the claim is bounded by the conditions tested, and every rebuttal below is an argument about whether those conditions are the ones nature hands you. The claim itself, that there exist regimes where parsimony wins, is not in dispute, and your slider finds one.
- The model rebuttal. Spencer, Susko & Roger, MBE 22:1161 (2005), argued the result is limited to a special case, that the deterministic assignment of the first half of sites to one regime and the second half to the other is not the right statistical model of the situation. Their abstract states that the earlier mixture model "was inconsistent because it was incorrectly implemented", and their answer is a correctly implemented mixture. That is what the purple curve is. We could not read the paper's body text (paywalled), so we do not quote any figure from it here.
- The parameter-range rebuttal. Gadagkar & Kumar, MBE 22:2139 (2005), re-ran the comparison over the full range of proportions of sites involved in heterotachy rather than the single 50:50 split, and concluded that likelihood is significantly superior overall even under heterotachy. That is what happens to the blue curve as you slide ρ away from 0.5.
- The design rebuttal. Philippe and colleagues, BMC Evol. Biol. 5:50 (2005), argued that the specific design simultaneously changed the level of heterotachy and the average terminal branch length, so the two effects are not separable in the published results, and that keeping all terminal branches equal is biologically unrealistic.
We are not adjudicating this. Our bench reproduces the shape of the disagreement rather than any paper's exact parameter grid: our regime assignment is random per site with probability ρ, not a fixed first-half/second-half split, and our branch-length values are our own. What the bench does establish, and what nobody in the exchange disputes, is the structural claim: whether parsimony or likelihood is more accurate here is a function of a nuisance parameter nobody can measure directly, and both sides were arguing about where in that parameter's range real data live. Slide ρ and you are holding the axis of the argument in your hand.
The check, and every free choice in it
What is exact, with no simulation involved
The infinite-data verdict, the three informative-pattern probabilities, the expected tree lengths and the zone boundary are not estimated. They are computed by summing the Jukes-Cantor pattern probability over both internal-node states for all 256 possible site patterns. At your current settings:
The boundary is found by bisection on the condition f(AC|BD) > f(AB|CD) to 200 iterations. At short = internal = 0.05 it sits at long = 0.394048.
What is simulated, and how to reproduce it byte for byte
Sequences are evolved site by site with a seeded mulberry32 generator. The curve at sequence length n uses seed 20260720 XOR n; the heterotachy bench at fraction ρ uses seed 7314159 XOR round(1000ρ). Given the same seeds, the same replicate counts and the same code, the numbers below are deterministic, and the verifier reproduces them exactly rather than to a tolerance.
The free choices we made, which change the numbers
- Jukes-Cantor for both generation and inference in layer 1. Equal base frequencies, one substitution rate. The inconsistency result is not an artefact of this choice (Felsenstein's original demonstration used a different two-state model), but the exact boundary 0.394048 belongs to Jukes-Cantor and to these branch lengths.
- Ties count as failures. When two topologies score equally we record it as not recovering the true tree. With short alignments this costs the true tree a few percent; it never changes the direction of any curve.
- Replicate counts of 120 (layer 1), 40 (the likelihood table) and 60 (the heterotachy bench). These set the width of the Wilson bars, which are drawn, not hidden.
- The ML optimiser is coordinate ascent with golden-section line searches, 3 sweeps over the 5 branch lengths at 14 iterations each (12 for the mixture model, over 11 parameters). This is a local search on a surface that can have multiple optima, so an occasional replicate may be scored on a sub-optimal fit. The verifier ships that control: it re-runs the whole heterotachy bench on identical seeds at double the sweeps and double the golden iterations, and the largest movement in any of the fifteen figures is 0.0167, one replicate out of sixty.
- The heterotachy design is ours, not a reproduction of any paper's: regime 1 puts A and C on long branches, regime 2 puts B and D on them, each site drawn independently with probability ρ. Long 0.75, short 0.10, internal 0.20, 1000 sites.
- The mixture ML model is given the right number of classes (two). Real analyses have to choose that number from the data. Handing it over is generous to the purple curve, and it is the generosity that makes the point: even fully informed, the fix is a model change, not a method change.
What we are not claiming
- Not that phylogenetics is unreliable. The precise claim is that consistency is conditional on the model containing the truth, and that the condition is checkable in simulation and not checkable in the field.
- Not that real analyses use Jukes-Cantor with homogeneous rates. They have not for decades: rate variation across sites, partitioned and mixture models, and model selection are standard. The demonstration is about what happens when a model is wrong in a way you did not anticipate, which is the only way models are ever wrong.
- Not that parsimony is better than likelihood, or worse. On this bench the ordering reverses with a single parameter, which is the finding.
- Not that Kolaczkowski & Thornton were right, or that their critics were. The exchange is presented with its rebuttals, and the bench shows the parameter it turned on.
- No novelty is claimed anywhere on this page. Everything here has been in the literature since 1978 or 2005.
Run it yourself: node research/more-data-wrong-tree/verify-more-data-wrong-tree.mjs. It recomputes every figure above from scratch, offline, with no dependencies, prints 115 checks, and exits nonzero if any of them has drifted. Its last section slices this page's own inline engine out of the HTML, executes it with no browser present, and runs it head to head against the reference implementation: same random stream, same alignment, same log-likelihoods to the last bit. So the page and the checker are not two descriptions of one program, they are two programs that have to agree.
What is idealised, and what is exactly true
Exactly true. For four taxa, only the 36 site patterns of the form "two taxa share one base, the other two share another" can separate the three unrooted topologies; the other 220 patterns cost the same on all three. So parsimony's answer is decided entirely by which of three probabilities is largest, and as sequence length grows the sample proportions converge on those probabilities almost surely. If the largest one belongs to a wrong topology, parsimony converges on that wrong topology with probability approaching 1. There is no sample size that rescues it, and no resampling scheme either.
Idealised. Four taxa, no insertions or deletions, no alignment error, independent sites, stationary and reversible substitution, and a substitution model (Jukes-Cantor) simpler than any model you would choose for a real dataset today. (JC69 has not vanished: it is still the null model inside model-selection procedures and still a selectable option everywhere. It is not what anyone selects.) Real long-branch attraction is fought with denser taxon sampling to break long branches, with better models, and with outgroup choice; none of that is in this bench.
About the word "wrong". The reconstructed tree AC|BD is not a bad estimate of AB|CD. It is a different tree. The failure is not imprecision that shrinks with data, it is bias that hardens with data, and that distinction is the whole reason consistency is the property people care about.
Numbers you may see move. The likelihood figures depend on the optimiser settings named in the check panel. The exact probabilities do not depend on anything except the branch lengths and the model, and are stable to the last digit shown.