The number turns on a software default
The Sample Nobody Finished
Chao and Jost standardised two Costa Rican beetle samples by completeness and reported that old growth is 4.7 times richer. That number needs a base coverage. Their own 2014 guidance and their own reference software pick different ones, and the headline moves by 1.66x.
A 237-individual sweep sample holds 112 observed morphospecies. Eight standard estimators turn that same spectrum into answers as high as 489. This is a methods demonstration, not a controlled conservation comparison: the two sites differ in effort, vegetation, and quite possibly in season and time of day.
loading
Loading the embedded spectrum.
The rule that sets the headline is not the rule in the paper
Coverage standardisation compares two samples at the same estimated completeness. Before it can do that, something has to choose which completeness. That choice is a pure convention with no biological content, and at least five versions of it are in print, three of them by the same group in three years.
| Rule and source | Base coverage | Old / second individuals | Ratio |
|---|
iNEXT and iNEXT.3D, at C=0.7212
4.64xThe rule printed in Chao and Jost 2012, and the one both packages implement.
1.66x
apart
Chao et al. 2014 Box 1(b), at C=0.9283
2.80xThe rule prescribed two years later by an overlapping author group.
Same two samples, same equations, same authors. A reader who follows the printed 2014 guidance and a reader who runs the group's own R package get headlines that differ by a factor of 1.66, and neither has made a mistake. The 2014 rule also asks for the old-growth arm to be extrapolated to 7.7x its own sample size, which Step 0 of the same box advises against for richness.
What this page did not do
The package behaviour was established by reading source, not by executing it. R was never run here. The claim rests on iNEXT/R/invChat.R taking min over Chat.Ind(x, 2*sum(x)), on the base == 'coverage' branch of iNEXT.3D/R/CommonFun.R doing the same, and on that package's own documentation describing the minimum coverage at double the reference sizes. Anyone with R should run iNEXT once on these two frequency vectors and check. Whether Box 1(b)'s authors would apply their rule to a pair of samples this unequal is also not something we can know from the printed box.
One curve, not two methods
The disagreement between the coverage-standardised 4.7x and the asymptotic 1.6x is usually presented as two methods that will not reconcile. They reconcile exactly. Equation 7 tends to S_obs + f0, which is Chao1, as the extrapolation grows without bound, and Equation 9a drives the coverage estimate to 1 over the same limit. The paper says so on page 2539. So the point at coverage 1 is the Chao1 ratio, and it is the terminal point of the coverage curve itself.
Green line: computed live from the two shipped spectra with Equations 5, 7, 4b and 9a. Gold dots: the 21 ratios printed in Table F3. Hover the plot for values.
The curve rises to 4.86x at coverage 0.780, then comes back down and converges on 1.6316x. Everything that looked like a fight between methods is a question about where on this one axis you stand, and the answer to that is fixed by convention, not by evidence.
The hump has a mechanism. At the published base coverage the old-growth arm holds 42.2% of its own estimated total while the second-growth arm holds 14.6%. Equal coverage in individuals is not equal completeness in species: 35.4% of old-growth individuals are singletons against 7.2% in second growth, so the two assemblages buy coverage at very different rates. At intermediate coverage the ratio is a joint function of richness and of the shape of the abundance distribution, and here the shape term dominates.
Which puts a number on the headline that the paper does not print: the published coverage-standardised ratio is 2.88 times the same machinery's own asymptote, on the same two samples, using the same equation. That is not a criticism of the arithmetic. It is a measure of how much of the famous 4.7x is a statement about abundance shape rather than about richness.
One more claim that inverts on the paper's own example
The abstract says coverage-based rarefaction "throws away less data than traditional size-based rarefaction". Measured as the fraction of the larger sample actually used, on this dataset it is the other way round.
| Comparison | Second-growth individuals used |
|---|---|
| size-based, m=237 | 24.3% |
| size-based, m=500 | 51.2% |
| coverage-based, no extrapolation | 5.6% |
| coverage-based, as printed | 9.2% |
The claim is data-dependent, and it holds when the sparse sample happens to have high coverage. It fails when the sparse sample has low coverage, which is exactly the case Example 1 was chosen to showcase. This reading takes "throws away data" to mean individuals discarded from the larger sample; the authors also say that with extrapolation no data need be discarded, and under that reading the counterexample does not land.
This framing may be a rediscovery
The identity above is arithmetic and it is checked below. What is not established is that it is new. Chao et al. 2020 (Ecol Res 35:292-314) frames the asymptotic-versus-standardised split directly, and Roswell, Dushoff and Winfree 2021 (Oikos 130:321-338) sits on the same question. Neither has been read here. If either states this, the page has reinvented it rather than found it.
| Coverage | Old m | Second m | Re-divided | As printed |
|---|
First, earn the anchor
The 2012 appendix prints an expected 70.19 species when the 976-individual second-growth sample is rarefied to 237 individuals. Equation 5, applied to the published frequency counts, gives:
Paper prints 70.19 · this page recomputes
Same rows, Hurlbert-Sanders rarefaction, no fitted parameter.
70.18934594
Rounded to the paper's two decimals, it is exactly 70.19. The unrounded residual against the printed value is -0.00065406 species.
The residual in the extrapolated column, diagnosed
The extrapolated old-growth cells do not reproduce exactly, and inverting Equation 7 against each printed value says why. Solving for the undetected-species count that would produce each cell gives a stable 351.97 to 352.07 across the six largest extrapolated sizes. But Equation 8 in the paper's own main text gives 351.31, and Appendix C of the same paper uses Chao 1984's 352.80. The table was produced by neither. The consequence is under 0.04% and of no scientific importance; it is an inconsistency between a paper's stated method and its published table, and the only way to see it is to run the arithmetic.
| Old-growth target size | Printed richness | Undetected species implied |
|---|
The two smallest extrapolated rows drift because Equation 7 is least sensitive to the undetected count close to the reference sample, so a value printed to two decimals pins it least well there. The whole inversion runs against values already rounded to two decimals, so a rounded intermediate inside the authors' own code cannot be excluded, and nothing on this page is built on the 352.00.
Now cross every choice
One grid asks two different questions. The first is how many species the old-growth assemblage might contain. The second is which site's richness is larger under a named standardisation rule. Keep those questions separate.
Choose one defensible specification
Loading the embedded spectra and running 6,912 cells.
Old-growth / second-growth richness
Waiting for the grid.
loading
A disabled control is inert for the selected basis. No result changes when an inert choice changes, but the full factorial grid retains those cells so its declared size remains 4 × 8 × 6 × 3 × 4 × 3.
Two of those four numbers are not evidence
The minimum and the maximum are properties of the data. The median and the share below parity are properties of how many variants were typed into each menu. This grid carries three singleton treatments, two of which push every cell below parity, and eight estimators rather than sixteen. Change those counts and the median moves without a single beetle changing. A specification curve is a range and a set of sign flips; its central tendency is a statement about the author's menu, and this page's median is quoted only so it can be labelled as one.
The plot reports any downsampling after it is drawn.
The published equal-size result, 112 divided by 70.1893, is 1.596x. One fully identified representative of that inertly duplicated specification sits at the 86.5% percentile of the full defined curve.
Scope changes the headline
With the published singletons left as recorded, the grid spans 0.80x to 4.77x. Across all singleton treatments it spans 0.22x to 4.77x. The 2016 correction estimates 26.67 old-growth singletons but 89.53 second-growth singletons, even though only 70 were recorded there. Every corrected-singleton comparison then favours second growth. That method was designed for spurious sequencing reads, not these 1970s morphospecies, so its leverage is a warning about transportability, not a verdict about either forest.
Four bases, four questions
The grid does not pretend raw counts, equal effort, equal completeness, and the entire assemblage are aliases. They estimate different objects.
| Basis | Minimum | Median | Maximum | Refused |
|---|
The asymptote is a family, and the grid holds part of it
Same recorded spectrum, ACE at k=10. These estimate a whole assemblage rather than equal fractions of two assemblages. The last three rows are computed here but are not in the grid dimension above, which is what a menu looks like from outside.
| Estimator | Old | Second | Ratio |
|---|
The eight in-grid options span 0.86x to 1.65x. The three higher-order jackknives, which are in EstimateS and in every comparison paper, give 1.11x to 1.35x. Brose, Martinez and Williams (2003) are cited as recommending jackknife orders 3 to 5 at the completeness this dataset has; that recommendation is second-hand here and has not been read at source, so it is offered as a reason those rows exist, not as an endorsement.
Which choice moved the answer?
On log ratios in the defined full grid, singletons carries the largest marginal share. This decomposition is indicative, not exact, because 864 capped extrapolations are refused and the surviving design is unbalanced. Inert factors can be diluted by repeated cells, and the cap's apparent share reflects selection among survivors rather than a changed formula.
The interaction remainder is 15.6%. Interactions matter here because singleton treatment changes sample size, coverage, and every estimator at once.
What the curve cannot see
This is the part that indicts the instrument above. Rank every degree of freedom in this problem by how far it moves the answer, and the third and fourth largest are not on the curve at all, because neither is an operation you can perform on a frequency vector.
Which two of Janzen's samples you hold
Chao and Jost compare one old-growth sample with one second-growth sample. Janzen took sweep samples at about 25 sites, with day and night, wet and dry, an elevational transect and repeat years, so there are far more than 25 samples. Four more Osa beetle spectra are recoverable from the published literature: two old-growth in Chao and Shen 2003 Table 5, two second-growth in Huillet and Paroissin, arXiv:0809.4181, Tables 7 and 8. All four transcriptions reproduce their printed species and individual totals exactly. That gives three old-growth and three second-growth samples, so nine pairings.
| Pairing | Equal effort | Equal n | Coverage, no extrap. | Coverage, 2012 rule | Chao1 |
|---|---|---|---|---|---|
| O1 x S1 (published) | 0.800x | 1.596x | 3.627x | 4.642x | 1.632x |
| O1 x S2 | 0.742x | 1.320x | 2.398x | 3.105x | 2.028x |
| O1 x S3 | 0.783x | 1.590x | 4.096x | 4.838x | 1.044x |
| O2 x S1 | 0.557x | 1.540x | 3.572x | 3.988x | 0.950x |
| O2 x S2 | 0.517x | 1.298x | 2.389x | 2.638x | 1.181x |
| O2 x S3 | 0.545x | 1.606x | 4.601x | 4.439x | 0.608x |
| O3 x S1 | 0.564x | 1.334x | 2.348x | 2.702x | 0.888x |
| O3 x S2 | 0.523x | 1.112x | 1.555x | 1.828x | 1.104x |
| O3 x S3 | 0.552x | 1.364x | 2.580x | 2.674x | 0.568x |
Across all nine pairings and all four standardisations the answer spans 0.517x to 4.838x. On the asymptotic comparison four of the nine pairings put the second-growth site ahead; at equal field effort all nine do. The published pair sits second of nine on both headline measures. That is not an allegation of cherry-picking, and there is no evidence the authors had the others to choose from. It is the observation that a lever worth about three-fold sits entirely outside the analytical frame, and that the realised choice sits high inside it.
Those four spectra are cited and recomputed rather than shipped. They are transcribed from a subscription journal table and an arXiv preprint, and this project does not redistribute data whose licence it has not established. The nine-pairing arithmetic runs in this page's verifier, against the same code the page uses, and each transcription is checked against its printed species and individual counts before anything is computed from it.
Whether two morphospecies are one species
Janzen measured his own error rate and published it. In Rev. Biol. Trop. 24(1):149-161, page 150, describing the same protocol: 7% of the 320 beetle morphospecies in the reference collections were found to be synonyms, with sexual dimorphism responsible for most of it, and with more attention the figure could have been reduced to about 3%. Sexual dimorphism splits one species into two, and the split halves land preferentially in the rare tail, which is exactly the singleton count Chao1 squares.
| Synonymy rate | Species after merging | Base coverage | Coverage-standardised | Asymptotic Chao1 | Observed count |
|---|
At the rate the collector himself reported, the coverage-standardised headline falls from 4.64x to 2.54x. The effect compounds, because merging singletons also raises the estimated coverage and pushes the base coverage over the top of the hump. This is not a sensitivity analysis over an invented perturbation; it is the published measurement uncertainty of the input, propagated.
The proxy is ours, and the point survives it
Merging morphospecies is not a frequency-vector operation. Doing it properly needs the specimens, and Janzen's specimen-level records are not published. What the table above does is merge singletons in pairs at Janzen's rate, which is our construction chosen for transparency, not a published correction. If the synonyms were spread evenly across abundance classes instead of concentrated in the rare tail, the effect would be smaller. The rate is Janzen's measurement; the mechanism is an assumption. Both of these levers, sample choice and species delimitation, are invisible to any curve built from a frequency vector, and a specification curve that looks exhaustive while omitting its two largest non-estimator levers is the exact failure this page is about.
Three settings a reader can reach where the comparison says nothing
- The bottom of any coverage or size control is an identity. Equation 5 at m=1 reduces to exactly 1 species for every assemblage: old growth 1.000000, second growth 1.000000. Any two assemblages are therefore exactly equally rich at one individual, with a standard error of exactly zero. The paper's own Table F3 prints this as its first row: 1.00 (0.00). A slider that reaches this region manufactures perfect agreement with perfect certainty.
- A sign flip at low coverage that is an artefact of integer sample sizes. Matching coverage 0.10 needs 2.82 and 3.16 individuals. Interpolating on m the ratio there is 0.8918; rounding each arm up to whole individuals it is 0.7580, which looks like a clean reversal of the paper's conclusion. At the same whole number of individuals in both arms it is 0.9944, which is the honest reading, and it is also what the paper's own Table F3 does at that row. The whole region below about coverage 0.2 compares three individuals with four.
- The default rule pins one arm to its own reference sample. At the no-extrapolation base coverage the old-growth arm is standardised to exactly its own 237 individuals, giving 112.0000 species. The classical conditional rarefaction variance (Heck, van Belle and Simberloff 1975) is exactly zero at m=n, so a page using it would collapse the interval on the arm carrying almost all of the uncertainty, precisely at the level the default rule selects. Colwell et al. 2012's unconditional variance exists to fix this. At the Chao et al. 2014 base coverage the same trap moves to the other arm: the second-growth sample is pinned at 976 individuals and the entire comparison becomes a statement about one extrapolated number.
Neither variance formula is implemented here, so the third item is an algebraic statement about Equation 5 at m=n rather than a computed one. The zero-variance claim in the first item is not: rarefaction to one individual is deterministic, so its bootstrap standard error is exactly zero by construction.
Five checks on this page cannot fail
Every one of these is an algebraic identity. Passing them is not evidence of anything, and a page that prints them as green lights has told the reader that it verified something when it did not.
- S(m=n) = S_obs, because the Equation 5 sum is empty at m=n.
- S(n+0) = S_obs, Equation 7 with no extrapolation.
- C(n+0) = C(n), Equation 9a with no extrapolation, which is Equation 4a.
- The rarefaction and extrapolation curves meet at the reference point, which follows from the two above. Colwell et al. 2012 list this meeting among their Important Findings and call it surprising. The values meeting is an identity.
- S(n+infinity) = Chao1, definitional in Equation 7.
What is not an identity, and is a real result, is the two curves' slopes meeting. Equation 4b and Equation 5 are separate printed formulae, so their agreement is a genuine cross-implementation check, and Equation 10 extends it past the reference point.
Equation 6, on the rarefaction branch
Worst absolute gap between the one-step richness increment and the coverage deficit, over eight sample sizes per site.
old growth: 3e-14
second growth: 1e-13
Equation 10, on the extrapolated branch
The same identity past the reference point holds only when Equation 7 is fed the Equation 8 undetected-species estimate. Seven of the eight menu options break it, which is how a check that can fail earns its keep.
| f0 estimate | Worst gap | Eq. 10 |
|---|
The published bootstrap recipe does not produce the published error bars
Appendix C says, verbatim, that from the bootstrap community "a random sample of m individuals is generated with replacement", and that Equation 5 is then applied to it. Implemented exactly as written, that gives standard errors far wider than the ones printed in the same paper. Drawing n individuals and then using Equation 5 to rarefy that draw down to m reproduces them.
| Cell | Printed s.e. | Appendix C read literally | Draw n, then rarefy |
|---|---|---|---|
| old growth, m=100 | 2.77 | 4.46 | 2.69 |
| second growth, m=89 | 1.02 | 3.58 | 0.94 |
Running the Appendix C bootstrap.
The documented procedure and the implemented procedure are not the same procedure. Anyone rebuilding the uncertainty bands from the appendix rather than from the package gets intervals up to three times too wide at small m, and nothing in the paper would tell them.
The check · live, including the awkward parts
Table F1 identities
Old: S=112, n=237
Second: S=140, n=976
All four totals match the appendix caption. A mistranscribed row would break this one.
Rarefaction anchor
27 / 27 printed richness cells match to two decimals.
70.19 recomputes as 70.18934594.
Reference coverage
Old: 0.6459 → printed 65%
Second: 0.9283 → printed 93%
Eq. 4b discrepancy resolved
Direct Eq. 4b matches the small-m printed cells after rounding.
old m=5: 0.154596; m=10: 0.236314
second m=5: 0.150618; m=10: 0.264744
The shifted ratio gives old 0.131563/0.223305 and second 0.123792/0.244316. That was the source of the scout's open discrepancy.
An external value caught what four internal ones missed
A published SpadeR run on this same old-growth sample prints iChao1 489.1. Building the Chiu et al. 2014 correction on the bias-corrected Chao1 gives 453.0; building it on Chao1 times (n-1)/n gives 488.78, which is the figure this page reports.
Every identity check on this page passes under both versions, because none of them touches f3 or f4. Only a value from outside could separate them, and the 0.32 that remains is unexplained.
What is not claimed
ACE-1 is omitted because its formula variant was not independently verified.
No permutation p-value is reported. Anonymous frequency spectra lack cross-site species identities, and two uncontrolled illustration sites do not support a meaningful site-label shuffle.
No Hill number above q=0 is computed here. The claim that the whole inflation is a richness-corner effect needs q=1 and q=2, and this page does not have them.
Why 864 specifications are refused
- Waiting for the live grid.
A reader can reach these cells. They display “not defined” and the selected cap reason rather than reusing the last valid number.
Corrections, data limits, and provenance
The ESA corrections log remains unchecked: both candidate log URLs returned HTTP 404. A targeted search found no correction, but a negative search is not a verified clean record.
The shipped data are the complete Table F1 frequency spectra, 35 abundance/count pairs. Species identities and specimen-level rows were not published in this appendix, so incidence models and a valid specimen-label permutation are outside this page's data scope.
Two published transcriptions of the second-growth sample disagree. Huillet and Paroissin's Table 6 prints an abundance class of 29 where Chao and Jost and Colwell et al. both print 19, which is exactly the 20-individual difference between the 996 that appears in some secondary literature and the 976 that four independent tables print. The row also breaks ascending order at that entry. 976 is right. All four of those tables may still descend from one original transcription of Janzen 1973a, which nobody in this chain appears to have read directly.
What this page cannot settle
- Nobody labels the season or the time of day of the old-growth sample. No source reached here describes the 237-individual sample beyond "Osa primary" or "Osa old-growth", and it is demonstrably not either of the dry-season 1967 primary-hill samples that Chao and Shen used. Janzen 1973b's entire subject is that season, time of day, elevation and insularity move these counts enormously. So it cannot be ruled out that the published old-growth against second-growth contrast is partly a day against night, or a wet against dry, contrast. Of everything unresolved here this matters most, because it is a confound in the headline itself rather than a choice about how to analyse it.
- The taxon. Janzen tabulated beetles, bugs and all arthropods, and reported that bugs respond to this vegetation contrast in the opposite direction. His synonymy rate was 7% for beetles and 1% for bugs, so the species-delimitation lever is taxon-specific too. No bug frequency vectors could be found in print.
- The sampling model. Each 800-sweep sample was taken in eight 100-sweep subsamples, which is exactly the structure a sample-based incidence analysis needs. Those subsample counts are not published. Colwell et al. 2012 state that intraspecific aggregation makes individual-based methods overestimate richness at reduced m, and foliage insects are strongly aggregated, so the direction of the bias is known and its net effect on a ratio with one arm rarefied and one extrapolated is not.
- Which question is the right one. Chao and Jost's case for coverage standardisation rests on the replication principle. Alroy 2020 asserts a different invariance, split-analyze-and-sum, and observes that coverage standardisation fails it while asymptotic extrapolators pass. Neither axiom implies the other and each camp's method fails the other's test. There is no neutral ground, and a page that supplies one has invented it.