The Verification Venue · a method that fails harder the more you feed it

More Data, More Certain, Still Wrong

Here is a four-species tree. We know it is the truth, because we are about to grow the DNA on it ourselves. Then we hand the DNA to a reconstruction method and ask for the tree back. In one region of branch-length space the method returns a tree that never existed, and returns it harder the more sequence you give it.

Every phylogeny you have ever seen is an inference. Nobody watched the branchings. What makes the inference trustworthy is supposed to be a guarantee called statistical consistency: as the data grow without bound, the estimate converges on the truth. In 1978 Joe Felsenstein showed that maximum parsimony, then the dominant method, does not have that guarantee. Worse than not having it: there are trees for which parsimony converges on a specific wrong answer, with the support for that wrong answer growing steadily as you add sites.

The bench below is that proof, made operable. The tree on the left is real by construction. Drag its branch lengths. The verdict panel computes, exactly and instantly, what parsimony would say if you handed it infinitely many sites.

1 · the true tree, and the infinite-data verdict in the zone

Units are expected substitutions per site. Two long branches that are not neighbours, separated by a short internal branch: that is the whole recipe. Presets:

With unlimited data, parsimony converges on

computing

Parsimony's whole decision rests on three numbers: the probability that a site shows two taxa with one base and two with another. Only 36 of the 256 possible site patterns separate the three topologies, and these are they. The bars are exact probabilities, not samples.

Zone boundary, this short/internal

·

long branches past this length flip the verdict

Expected tree length, true topology

·

steps per site, AB|CD

Expected tree length, AC|BD

·

parsimony picks the smaller

The map below is the whole parameter plane, coloured by that same exact calculation: which topology wins the informative-pattern count when the data are unlimited. The pale region is safe. The dark region is what the literature calls the Felsenstein zone. Your current settings are the ring.

the zone, computed cell by cell internal branch fixed at 0.05
parsimony converges on the true tree parsimony converges on AC|BD your settings

Now grow the data

The exact verdict says what happens at infinity. Here is what happens on the way there. For each sequence length we simulate whole alignments down the true tree under Jukes-Cantor, score all three topologies by Fitch parsimony, and record how often the true one wins outright. The blue curve is a safe tree (long branches 0.25). The red curve is the same tree with one number changed: the long branches lengthened to 0.45.

2 · accuracy against sequence length
safe tree, long = 0.25 your tree, long = 0.45 1/3, blind guess bars: 95% Wilson intervals

True tree recovered at 100 sites

·

of 120 replicates

At 10 000 sites

·

of 120 replicates

Change

·

a hundredfold more data

Read the red curve again, slowly. It is not noisy, and it is not flat. It goes down. A hundredfold increase in evidence takes the method from occasionally right to reliably wrong, and it would keep going: the wrong topology's informative-pattern probability exceeds the true one's by a fixed margin, so the error is a law of large numbers working perfectly, on the wrong quantity. Bootstrap support for the wrong tree climbs toward 100% for exactly the same reason, resampling the sites cannot escape what the sites say. That last point is not computed on this bench; it is a known result, shown by simulation in Hillis & Bull, Syst. Biol. 42:182-192 (1993), and it follows from the bootstrap proportion tracking the same asymptotic winner. Confidence measures repeatability, not truth.

The mechanism has a name, long-branch attraction. Two lineages that have each changed a great deal will, by chance alone, sometimes land on the same base at the same site. Parsimony reads shared bases as shared ancestry, and cannot tell a coincidence from an inheritance. The two long branches are given many chances to coincide; the short internal branch, which carries the real signal, is given almost none.

This is not only a simulation result. In the 1990s microsporidia (fast-evolving parasites with reduced genomes) were placed at the base of the eukaryote tree by early rRNA analyses, and were later reassigned to the fungi once the long-branch artefact was recognised and denser taxon sampling and better models were applied. The Ecdysozoa hypothesis, that arthropods and nematodes share a moulting clade, was resisted in part because the fast-evolving nematode Caenorhabditis was being pulled toward other long branches. Both cases are now settled; they are here as history, not as open questions.

The answer everyone gives

If you have taken a phylogenetics course you already know the reply, and it is correct: parsimony is naive, use maximum likelihood, which is consistent. So let us do that, on exactly the data that just broke parsimony. Same simulation, same zone, but the three topologies are now scored by fitting Jukes-Cantor branch lengths and comparing log-likelihoods.

3 · likelihood, with the true model, in the same zone
sitesparsimonylikelihood (JC)

Same seeds, same alignments, same branch lengths (long 0.45, short 0.05, internal 0.05). Forty replicates per row.

Likelihood wins, decisively, and it wins because more data helps it. That is consistency, demonstrated rather than asserted. The lecture could end here, and usually does.

Where consistency actually lives

It should not end here, because the guarantee has a precondition that gets dropped in the retelling. Likelihood is consistent when the model contains the truth. The theorem is about the pair (method, model), and the word "consistent" attaches to the pair. Take the truth out of the model and the guarantee goes with it.

So here is a process the Jukes-Cantor tree model does not contain. Real sites do not all evolve at the same relative rate down every lineage; the pattern changes over time, a phenomenon called heterotachy. We build the simplest possible version: the true topology stays AB|CD throughout, but a fraction ρ of sites evolve with A and C on long branches, and the remaining sites evolve with B and D on the long branches instead. Every site is generated on the true tree. Nothing is faked. The alignment simply has no single set of branch lengths that describes all of it.

Three methods now compete on that alignment: parsimony, likelihood fitting one homogeneous Jukes-Cantor tree (the misspecified model), and likelihood fitting a two-class mixture of branch lengths with the class weight free (a model that does contain the generating process). Slide ρ and run.

4 · the heterotachy bench

ρ = 0 or 1 is a plain homogeneous tree. ρ = 0.5 is the 50:50 split the published exchange was fought over.

Short branches 0.10, internal branch 0.20, 1000 sites, 60 replicates.

maximum parsimony ML, one JC tree (misspecified) ML, 2-class mixture (contains the truth)

At ρ near 0 or 1 there is no heterotachy, the homogeneous model is correct, and likelihood is far ahead. As ρ approaches 0.5 the misspecification bites, and the blue curve falls below the green one: on this bench, at that setting, parsimony beats likelihood. Parsimony did not improve. Its curve is nearly flat across the whole range: it is not good here, it is merely unbothered, because it never made a claim about the process in the first place. What changed is that likelihood's claim became false.

And the purple curve is the point. Give likelihood a model that contains the heterogeneity and it is back on top wherever the heterogeneity is real, without changing the method at all. Read the endpoints honestly, though: at ρ = 0 and ρ = 1 there is no heterotachy to model, and the mixture pays for its extra parameters, finishing 0.0833 and 0.0333 below plain Jukes-Cantor likelihood on this bench. That is the ordinary cost of fitting a richer model than the data needs, and it is small next to the 0.8833 the plain model loses at ρ = 0.50 by being wrong. Consistency was never a property of "likelihood". It was a property of a model that happened to be right.

The part that is still an argument

That green-beats-blue crossing is not a curiosity we invented. It is the subject of a sharp published exchange, and the exchange is worth laying out because the disagreement is about exactly the quantity you were just sliding.

Evolutionary rates vary among sites and across the phylogenetic tree (heterotachy). A recent analysis suggested that parsimony can be better than standard likelihood at recovering the true tree given heterotachy. Spencer, Susko & Roger, Mol. Biol. Evol. 22:1161 (2005), opening lines

We are not adjudicating this. Our bench reproduces the shape of the disagreement rather than any paper's exact parameter grid: our regime assignment is random per site with probability ρ, not a fixed first-half/second-half split, and our branch-length values are our own. What the bench does establish, and what nobody in the exchange disputes, is the structural claim: whether parsimony or likelihood is more accurate here is a function of a nuisance parameter nobody can measure directly, and both sides were arguing about where in that parameter's range real data live. Slide ρ and you are holding the axis of the argument in your hand.

The check, and every free choice in it

What is exact, with no simulation involved

The infinite-data verdict, the three informative-pattern probabilities, the expected tree lengths and the zone boundary are not estimated. They are computed by summing the Jukes-Cantor pattern probability over both internal-node states for all 256 possible site patterns. At your current settings:

·

The boundary is found by bisection on the condition f(AC|BD) > f(AB|CD) to 200 iterations. At short = internal = 0.05 it sits at long = 0.394048.

What is simulated, and how to reproduce it byte for byte

Sequences are evolved site by site with a seeded mulberry32 generator. The curve at sequence length n uses seed 20260720 XOR n; the heterotachy bench at fraction ρ uses seed 7314159 XOR round(1000ρ). Given the same seeds, the same replicate counts and the same code, the numbers below are deterministic, and the verifier reproduces them exactly rather than to a tolerance.

·

The free choices we made, which change the numbers

What we are not claiming

Run it yourself: node research/more-data-wrong-tree/verify-more-data-wrong-tree.mjs. It recomputes every figure above from scratch, offline, with no dependencies, prints 115 checks, and exits nonzero if any of them has drifted. Its last section slices this page's own inline engine out of the HTML, executes it with no browser present, and runs it head to head against the reference implementation: same random stream, same alignment, same log-likelihoods to the last bit. So the page and the checker are not two descriptions of one program, they are two programs that have to agree.

What is idealised, and what is exactly true

Exactly true. For four taxa, only the 36 site patterns of the form "two taxa share one base, the other two share another" can separate the three unrooted topologies; the other 220 patterns cost the same on all three. So parsimony's answer is decided entirely by which of three probabilities is largest, and as sequence length grows the sample proportions converge on those probabilities almost surely. If the largest one belongs to a wrong topology, parsimony converges on that wrong topology with probability approaching 1. There is no sample size that rescues it, and no resampling scheme either.

Idealised. Four taxa, no insertions or deletions, no alignment error, independent sites, stationary and reversible substitution, and a substitution model (Jukes-Cantor) simpler than any model you would choose for a real dataset today. (JC69 has not vanished: it is still the null model inside model-selection procedures and still a selectable option everywhere. It is not what anyone selects.) Real long-branch attraction is fought with denser taxon sampling to break long branches, with better models, and with outgroup choice; none of that is in this bench.

About the word "wrong". The reconstructed tree AC|BD is not a bad estimate of AB|CD. It is a different tree. The failure is not imprecision that shrinks with data, it is bias that hardens with data, and that distinction is the whole reason consistency is the property people care about.

Numbers you may see move. The likelihood figures depend on the optimiser settings named in the check panel. The exact probabilities do not depend on anything except the branch lengths and the model, and are stable to the last digit shown.