One printed result · every defensible choice

The Corner Where the Number Turns Negative

A paper reports −0.1% growth when public debt reaches 90% of GDP. Recompute that point from the deposited panel, then watch all 1,620 defensible analyses arrive. The magnitude collapses. The direction mostly does not. And the choice everyone argued about moves the answer 1.63 points, in a way a grid of independent toggles is built to miss.

First, reproduce the published corner

paper: −0.1%loading rows

The NBER appendix prints −0.1%. Applying the spreadsheet range error, the selective-year rule, country weighting, and the separately recorded New Zealand value to the deposited rows recomputes −0.0621%, which rounds to the printed number. With all four corrections, the same top bin is loading, matching Herndon, Ash and Pollin's printed +2.2%.

Computing residual.

That is the anchor, not the conclusion. The question is what happens when reasonable choices move together: where the line goes, when debt is measured, which years enter, how observations count, and which summary sits in the middle.

The live specification

Move one choice. Watch the estimate move.

It opens at the historical 90% corner. The two red switches are errors, not defensible choices, so they remain outside the 1,620-cell grid. Every estimate carries the number of country-years and the number of countries behind it, and refuses to print at all when the high-debt side stops being a comparison between countries.

Loading the licensed panel.

Error, outside the grid Drops Australia, Austria, Belgium, Canada and Denmark. Three of them have 90%+ rows: 35 top-bin country-years in the corrected panel, or 25 at the historical corner, where the selective-year rule has already removed Australia's five and Canada's five.
Error, outside the grid The deposited 1951 growth value is −7.6351. This switch applies only at the historical country-weighted 90% corner.

Loading data/rr-panel.csv.

All choices crossed, errors off

The published magnitude occupies a thin tail

Computing the specification curve.

Preparing all curve columns.

Which sentence survives

Reinhart and Rogoff wrote several claims, not one

Each row is a reading of the published argument, checked against every cell of the grid with both errors switched off. The direction is robust, the size is not, and the two are routinely quoted as if they stood or fell together.

ReadingSpecificationsShare

Second layer · the choice that looks inert

The exclusion everyone argued about swings 1.63 points

Reinhart and Rogoff dropped Australia and Canada's 1946 to 1950 years and New Zealand's 1946 to 1949, and kept the United States' four bad demobilisation years at the same debt levels. Herndon, Ash and Pollin's central complaint is that asymmetry. On the crossed grid its main effect is almost nothing, which is the wrong answer and an instructive one.

Read the table across, not down. Under country-year weighting the exclusion costs a quarter of a point. Under the country weighting Reinhart and Rogoff actually used, it costs 1.63. The exclusion is not a small choice that happened to matter; it is a choice whose whole effect lives inside another choice.

Including this page's own grid

A grid of independent toggles reports it harmless

This is the failure the section exists to expose, and this page is not exempt from it. The 1,620-cell grid also crosses the threshold, the debt timing and the sample window. Averaging the weighting-by-exclusion interaction over all of those flattens it almost to zero.

So the page runs the kit's inertFactors() guard rather than reading main effects off a bar chart. A factor may only be called inert if it is inert on the whole grid and inside every fixed threshold, timing and window slice. The early-year rule fails the second test, and the guard forbids the claim.

The general lesson is the uncomfortable one. A specification curve of independent toggles measures each choice with the others averaged away, and any choice whose effect is conditional will be reported as no choice at all. On this dataset that choice happens to be the mechanism of the original scandal.

The fairest thing on this page

The published headline needs no error at all

loading

Computing the error-free route.

Weighting each contiguous spell above the line once, rather than each country or each country-year, is a defensible answer to the objection that nineteen consecutive British years are not nineteen independent facts. Reinhart, Reinhart and Rogoff adopted an episode framing themselves in Public Debt Overhangs: Advanced-Economy Episodes since 1800, Journal of Economic Perspectives 26(3), 2012, 69 to 86.

Combine it with a symmetric early-year rule, dropping 1946 to 1950 for every country instead of for three, and the corrected data reaches the published headline with no spreadsheet error and no transcription error in it. That is a sharper and fairer statement than saying the number required the mistakes, and it moves the argument onto the weighting question, which was never a mistake and never disclosed.

Why it lands there

Nine episodes, and the smallest one carries it

Dropping 1946 to 1950 for everyone leaves New Zealand 1951 stranded as a one-year episode at −7.6%, and episode weighting promotes it to a full unit of weight beside twenty-two Belgian years.

    That is the counterweight this route has to carry. Episode weighting reaches the published number only if a single year counts as an episode. Under the five-year definition its own authors later published, the same route gives a positive figure, and the page prints both.

    No threshold can create a cliff

    Look at the years before drawing the line

    The dots are country-years, not a fitted causal model. Move the threshold above. The line slides over a cloud whose local means rise again after the 90–100 band.

    Magnitude collapse, not a debunk

    The direction survives more often than the number

    91.0%

    In 1,474 of 1,620 specifications, growth on the higher-debt side of the selected line is lower than growth on its lower-debt side. The median gap is −0.72 percentage points.

    This is an association in this panel. It does not establish which direction causation runs. It does establish that “there is no relationship” would be a false summary.

    The number 90 was asserted, not estimated

    Move the line and nothing snaps

    Reinhart and Rogoff's own footnote says the brackets come from “our interpretation of much of the literature” and parallel the World Bank income groupings, and adds that a different set of cutoffs “merits exploration”. Here is that exploration, on their data, with country-year weighting and no exclusions.

    LineCountry-yearsCountriesAt or aboveBelowGap60% to line

    Three ways to make the comparison say nothing

    Settings that still print a number

    Every interactive instrument has corners where the arithmetic keeps working after the question has stopped meaning anything. These are this one's, found by looking for them.

    One: a window that leaves one country

    Narrow the sample period and raise the line and the “average growth of advanced economies above 90% debt” quietly becomes the average growth of Britain. Herndon, Ash and Pollin hit this themselves: their replication code carries the comment ## 1955-1980 has only Britain in the highest public debt category!!, and it never reached their paper. The sweep below crosses nine windows against six thresholds using the same estimator the controls use, and the same floor refuses the same cells.

    SettingCountry-yearsCountriesVerdictValue if printed

    Three of these windows are on the control above, grouped as narrow windows. They are offered so the refusal is reachable, and kept out of the crossing for two reasons: a decomposition needs a complete grid, and windows chosen because they degenerate would drag the whole curve with them.

    Two: a control that absorbs the treatment

    Add a linear debt/GDP term and ask whether debt above 90% still matters. The dummy and the level are nearly collinear over the range that matters, so the question answers itself and looks like evidence.

    SpecificationCoefficient on debt ≥ 90%Standard errorp

    Three: a normalisation that partly divides by itself

    Debt/GDP in year t carries nominal GDP of year t in its denominator, while the outcome is real growth into year t. A bad year mechanically raises the ratio and can push a country over the line. That guarantees some negative contemporaneous association before any economics happens, and it is measurable: compare debt observed before the growth year against debt observed after it.

    Debt measuredCountry-years above 90%Gap in growth

    Variance anatomy, with its own caveat

    What moves the answer, once each choice is averaged over the rest

    The threshold carries 24.8% of the grid's variance and the debt timing 23.2%, both of them choices nobody argued about. Weighting carries 0.3% as a one-way effect and the early-year rule 0.7%, and the section above is the reason those two numbers must not be read as importance.

      Interactions carry the remaining computing. The design is complete and balanced, so these shares partition the variance exactly. Exactly partitioned is not the same as correctly attributed.

      Three moves into the published corner

      The famous error is not the largest piece

      Across the complete 2×2×2 anatomy of the two errors and the one contested judgement, all seven effects are listed, main effects and interactions together, because separating them was how the interaction went missing in the first place.

        Herndon, Ash and Pollin Table 3

        All eight cells, from the same shipped rows

        “Printed” is the paper's one-decimal table. “Live” is recomputed in this browser. The New Zealand transcription is a ninth, separate step used only to reach −0.0621 and its printed −0.1.

        CombinationLivePrintedRounded

        An open question, not a finding

        Four countries share one growth history

        Scan the shipped panel for countries whose annual growth series are identical to each other and something turns up that we have not seen reported. It changes no conclusion in this dispute. We are printing it because it is in the file that a decade of argument rests on, and because we cannot tell you how it got there.

        What this is not

        It is not an accusation against anyone. Reinhart and Rogoff's working spreadsheet reached the public only as the simplified copy Herndon, Ash and Pollin published, and the processed panel this page ships descends from that same upload. We therefore cannot distinguish a copy error in the original workbook from one introduced when it was simplified for release, and we are not going to guess between two sets of authors on the strength of a file that passed through both.

        What would settle it

        Reinhart and Rogoff's own unmodified workbook, which we could not obtain. Short of that: our scan only catches series that match exactly across at least twenty overlapping years, so a partial or shifted copy would pass it unnoticed, and the number above is a floor rather than a count. The one route by which this could touch the headline runs through a shared denominator rather than a shared growth rate, and that route we can only describe, not test, because the panel we ship carries the finished ratio and not the GDP series underneath it.

        The verification venue

        The check

        Every result below is recomputed from the same shipped CSV that drives the controls. The verifier repeats these calculations under plain Node. Two of the checks are here because an earlier version of this page carried their weaker cousins: “no specification failed” was true only because nothing in the grid could fail, and “the early-year rule barely matters” was true only of a main effect.

        Anchor ladder

        Full grid, with its floor

        Degeneracy sweep, which finds failures

        The inert-claim guard

        Within-country null

        Counts

        What the check does and does not know

        Honest apparatus

        The data

        The page ships a four-column, 31,992-byte trim of RR-processed.dta from Zenodo record 4017423, deposited by Thomas Herndon, Michael Ash and Robert Pollin. It retains country, year, real GDP growth and central-government gross debt/GDP, drops rows missing growth or debt, sorts by country and year, and rounds the two measures to four decimals.

        Shipped panel · source and derivation · AGPL-3.0-only licence

        The biggest thing this grid cannot see

        Every cell above runs on one panel. An independent re-derivation that varied the panel found that choice moved the answer further than any analysis choice: the same corrected top bin is 2.17 on this data, 1.62 on Maddison per-capita growth, and 1.44 on either of two IMF debt series, against a spread of 0.33 for the next-largest factor. None of that is reachable from a control on this page, and no amount of crossing the choices we do offer would reveal it. Gross general government debt, the measure every actual austerity argument used, is not in this archive at all.

        The comparison

        Each grid cell compares growth at or above its selected line with growth below that same line. “Spell” means a maximal run of consecutive qualifying years within one country. Lagged debt comes from the same country and exact prior calendar year. A 10% trim removes floor(0.1n) values from each tail. The floor refuses any high-debt arm inside a single country or under five country-years, and flags any arm under ten.

        The formula we did not see

        The spreadsheet error is usually quoted as an averaging range stopping five rows early. We do not assert that range. The summary worksheet containing it is in no publicly available copy of the workbook, only the twenty country sheets were released, and every published statement of the range traces to one footnote rather than to a file. What reproduces here is the error's consequence: drop the five alphabetically first countries and the printed numbers appear. The range itself is a citation.

        The counts that tie, and the ones that do not

        The corrected panel has 110 observations at 90%+. The selective exclusion removes 14, leaving 96, exactly RR's printed top-bin count. Their own two counts of their own sample do not agree with each other: the body text says 1,186 annual observations and the note under Figure 2 says 1,180, with bucket counts of 443, 442, 199 and 96. This complete-case archive has 1,175. Neither gap is explained.

        The bins that disagree

        Our reconstruction of RR's choices gives 4.09 / 2.85 / 3.40 / −0.02. HAP print 4.1 / 2.9 / 3.4 / 0.0 for that reconstruction. RR's NBER appendix instead prints 4.1 / 2.8 / 2.8 / −0.1, so RR's own figure and RR's own appendix table disagree by 0.6 points in the 60 to 90 cell. All three rows are named instead of blended.

        Degrees of freedom and the choices excluded from the grid

        The grid is 3 weighting rules × 3 early-year rules × 3 central statistics × 5 thresholds × 3 debt timings × 4 sample windows. The spreadsheet range error and the New Zealand transcription are labelled switches outside it because neither is a defensible analysis choice. The three narrow windows on the control are swept rather than crossed.

        A net-debt option is absent because the archive contains no net-debt series. Only central-government gross debt is available. Weighting remains a live methodological dispute, which is why the page offers it as a control rather than deciding it in prose.

        What nothing on this page bears on

        Emerging markets and the 1790 to 2009 sample are separate claims with their own data and their own weighting, and neither is touched here. Neither is inflation, which shares the original figure. Population and GDP weighting are defensible and not implemented. No controls beyond the debt level and fixed effects are run: no investment, openness, initial income or banking-crisis terms, no instruments, no estimators that allow the relationship to differ by country. The endogenous threshold is a bootstrap of a point estimate and not Hansen's test of whether a threshold exists at all.

        Most of all, the 1,620 point estimates carry no inference bands. A specification curve without them looks more decisive than it is, and this is one. The clustered regressions in the vacuous-settings section are the only inference on the page, and the count of specifications agreeing with anything depends on how many variants we chose to enumerate, which is a decision and not a measurement.

        Sources and publication precision

        The digits −0.1 appear in Reinhart and Rogoff, Growth in a Time of Debt, NBER Working Paper 15639, Appendix Table 1. The published AER Papers & Proceedings article shows a bar; Herndon, Ash and Pollin describe their reading of it as approximate. We did not inspect that blocked publisher PDF.

        The corrected +2.2 and eight-cell ladder appear in Herndon, Ash and Pollin, Does High Public Debt Consistently Stifle Economic Growth?, Table 3. Their deposit carries the redistribution licence linked above. We worked from the April 2013 working paper and not the revised Cambridge Journal of Economics version.

        The episode framing is Reinhart, Reinhart and Rogoff, Public Debt Overhangs: Advanced-Economy Episodes since 1800, Journal of Economic Perspectives 26(3), 2012, 69 to 86. The reverse-causality comparison follows Arindrajit Dube's 2013 note on debt, growth and causality, re-derived here from the same rows.