Artificial Wasteland · cities, measurement, and the modifiable areal unit problem
A Million or So Exponents
Urban scaling says a city of two million produces more than twice what a city of one million produces, per head. Two published objections attack that on axes that never meet: where you draw the city, and what you assume about the noise. Neither paper runs the other's experiment. Here both are run on the same 3,080 counties and crossed. Then the exponent is handed to a zoning optimiser from 1979.
The claim is one of the most quoted results in the science of cities. Double a city's population and its patents, its wages, its GDP go up by more than double, by a fixed power. The power is the finding. This page is about how much of that power is in the cities and how much is in the map.
Everything below is computed from public files by code you can run: county GDP and wages from the Bureau of Economic Analysis, population from the Census Bureau, the county adjacency list, the official metropolitan delineation, and the county-to-county commuting matrix. The instrument in the next section is not a picture of a result. It is the same code, running in your browser, on the same 3,080 counties.
Draw the city yourself
A "city" here is a set of counties. The oldest way to build one from scratch, and the one Elsa Arcaute and colleagues used in 2015, is a density threshold: keep every unit above some number of people per square mile, glue the neighbours together, and call each connected blob a city. Drag the threshold. The map, the scatter and the exponent all move together, and the national totals at the bottom never move at all, because nothing has been added or taken away. Only the lines have moved.
the same cities, under five assumptions about the noise
Two things are worth doing before reading on. Put the threshold at the far left, where the whole eastern seaboard fuses into a single blob and there is no system of cities left to fit. Then put it somewhere sensible and switch the indicator to employment, where the exponent sits close to one wherever you put the line, and to retail trade, where it does not.
What the founding paper actually says
The reference is Bettencourt, Lobo, Helbing, Kühnert and West, Growth, innovation, scaling, and the pace of life in cities, PNAS 2007. Its Table 1 is twenty-three rows, and the first correction worth making is that there is no single exponent in it. The socioeconomic rows run from 1.07 to 1.34. The famous figure quoted as "1.15" appears in the table twice, as two individual rows, and is never asserted as the characteristic value. What the abstract asserts is β ≈ 1.2; the body says β ≈ 1.1–1.3; the supporting information derives 1.14 to 1.28 from a contact-counting argument and adds that "there is no strong indication that they must be identical for different urban systems." The sublinear road-surface row, the one usually cited for infrastructure economies, rests on twenty-nine German cities.
The second thing worth noticing is that the paper flags the exposure itself, in one sentence, and then does not return to it:
We adopted a definition of cities that is as much as possible devoid of arbitrary political or geographic boundaries, as integrated economic and social units, usually referred to as unified labor markets… In the U.S., these definitions correspond to metropolitan statistical areas (MSAs)… More detailed definitions of city boundaries are desirable and an active topic of research in urban geography.
Three of Table 1's rows are US metropolitan quantities the BEA still publishes, in years inside the window used here. Recomputing two of them is not an attempt to catch anyone out. It establishes that this pipeline lands where the literature lands when it is pointed at the same thing, so that when it lands somewhere else, the difference is the definition and not the plumbing.
Two objections that never met
Arcaute, Hatna, Ferguson, Youn, Johansson and Batty (2015) built systems of cities out of 8,850 England-and-Wales wards, sweeping a density threshold and then a commuting threshold, and produced more than twenty thousand different city systems from one dataset. Their finding: most indicators scale linearly regardless of the definition, and where nonlinearity does appear, the exponent "fluctuates considerably" with the definition. Their axis is the map.
Leitão, Miotto, Gerlach and Altmann (2016) attacked something else entirely. Fitting a straight line to log-log data is the maximum-likelihood estimator of exactly one noise model, and they wrote down five models that all satisfy the same mean relation and differ only in their fluctuations. Fitted to fifteen datasets, the models are rejected by the data in most cases, and they report that "in extreme cases, even the conclusion on whether a city index scales linearly or non-linearly with city population depends on the assumptions on the fluctuation." Their axis is the noise.
Both papers are right about their own axis. Neither runs the other's experiment, so neither of them says which one moves the answer more, and I have not found the comparison made elsewhere. That is a search, not a proof: the literature on urban scaling is large and a negative about all of it is not something this page can establish. The five models are implemented here from the equations in the paper: log-normal with the variance exponent fixed at 2 (which is ordinary least squares), the same with that exponent free, Gaussian with it fixed at 1, Gaussian with it free, and the multinomial person-model in which each unit of output is a token handed to a person. Every admissible city definition is fitted with all five.
every boundary definition × every fluctuation model
The interaction is the part neither paper could have seen alone. The fluctuation models do not disagree by a fixed amount: they agree almost exactly on the density-defined systems and diverge most on raw counties, because a model's answer depends on whether it is dominated by the many small units or the few large ones, and the boundary rule is precisely what decides how many small units there are. Leitão's paper predicts this in words. It reads, in its own discussion, as an explanation of why the "aggregation of cities (different city borders)" influences the estimated exponent at all. What it does not do is measure which effect is larger, because it never varies the borders.
What the threshold is really doing
A density threshold on English wards carves cities out of countryside. A density threshold on US counties does something coarser, because a county that contains a city also contains its farmland, and because the counties east of the Mississippi are small enough that at a low threshold they percolate: one connected blob from Boston to Richmond and inland to Chicago. The curve below is that transition. It is the reason definitions below about thirty people per square mile are not admitted to the comparison: there is no system of cities there to fit.
Arcaute chose her working threshold by exactly this kind of look: she watched the third-largest cluster and picked the value just before Liverpool and Manchester merged. The equivalent moment here is coarser and further down, and it is shown rather than described.
the commuting axis, and the minimum-size axis
Her second algorithm grows each city by commuting: a unit joins the city that takes the largest share of its commuters, if that share clears a threshold. Run over the 40 million Americans who commute across a county line, on the same ladder of density thresholds, it gives a second surface. A third knob, the minimum population a cluster needs to count as a city at all, is the one Leitão singles out as acting like a boundary change, and it does.
commuting threshold × density threshold
minimum city population
Iowa still has ninety-nine counties
None of this is new, and the part that is not new is older than the scaling literature by thirty years. In 1979 Stan Openshaw and Peter Taylor published A million or so correlation coefficients. They took the 99 counties of Iowa, correlated the percentage of the population over sixty with the percentage voting Republican, and then regrouped the counties into contiguous zones. The correlation went where they wanted it to go. Their restatement in 1983 is blunt about what the machinery is:
This can be regarded as an exercise in applied gerrymandering or, if you prefer, spatial engineering of zoning systems… for a 6 region aggregation of the 99 Iowa counties the range of possible correlations is between −.99 and +.99.
Their 1970 inputs sit behind ICPSR and NHGIS, both of which refuse an anonymous request, so their numbers cannot be recomputed and are quoted here, never presented as reproduced. The experiment, though, reproduces on data anyone can download: the same 99 counties, the share aged 60 and over from the 2023 Census estimates, and the Republican share of the 2020 presidential vote.
The exponent, gerrymandered
Here is the join. Openshaw's Table 12 is not about correlations at all. It is his zoning optimiser pointed at a regression slope, and it reports how far the slope moved: at six zones, from −121 to +27. A scaling exponent is a regression slope. And neither of the two critiques above cites him: the strings Openshaw, areal unit and MAUP appear nowhere in either preprint, which the verifier checks rather than assumes.
So: hold every county's population and output exactly as measured, partition all 3,080 of them into contiguous zones, and run his procedure on the exponent. The three maps below are the same data three times.
how far the exponent can be driven, by the aggregation alone
The obvious objection is that the optimiser is cheating with tiny zones, and it is worth putting the number on that rather than arguing about it. The unconstrained minimiser does contain a zone of 169 people. So the search was rerun with a population floor: below the floor a zone is merged into a neighbour before the search starts, and no move may take a zone under it afterwards.
And the second objection, which cuts the other way and is the more important one, Openshaw himself prints. He quotes Yule and Kendall, from 1950:
the student should not now go to the other extreme and claim that, since a large range of values of correlation coefficients may be obtained according to the choice of a modifiable unit, a particular value has no significance
That is right, and it is the sentence this page is built around. The zoning experiment does not show that the urban scaling exponent is meaningless. It shows how large the degree of freedom is that a city definition spends, and it comes with its own control: a contiguous zoning drawn at random, with no optimiser at all, lands in a narrow band. The freedom is enormous and almost none of it is used, because the conventions are narrow. Which is to say the stability of the exponent is a fact about the conventions at least as much as it is a fact about cities.
How far the answer travels
One indicator in one year would be a coincidence rather than a finding, so the whole grid is rerun across ten quantities and across every year from 2001 to 2019.
ten indicators, same 34 definitions
nineteen years, same 34 definitions
What this shows, and what it does not
It does not show that urban scaling is wrong. Across every admissible definition and every fluctuation model, the exponent for GDP stays in a band whose bottom is just below one and whose top is well above it, and superlinearity is the most common verdict by a wide margin. Output does concentrate. What moves is how much, and whether the movement clears the bar for calling the relationship nonlinear at all.
It does show that the two published objections are not the same size. On this dataset, for every one of ten indicators, redrawing the city moves the exponent further than changing the noise model does. The ratio is not constant, which is itself worth stating: it runs from a little over one to about twelve depending on the quantity.
And it shows the two are not independent. How much the noise model matters is set by the boundary rule, because the boundary rule sets how many small cities there are and the models differ mostly in how much attention they pay to small cities. A robustness check that varies one and holds the other fixed will under-report both.
the honest limits
- A county is a coarse atom. Arcaute's wards are a few square miles; the median county here is far larger. A coarser atom can express fewer distinct city systems, so the boundary range measured here is a lower bound on what her construction would give, not an upper one.
- The suppressed cells are not random. BEA withholds a county figure when it would disclose a single firm, and that hits small counties. Rather than drop those counties, which would move the boundaries between indicators, any city containing one is dropped for that indicator alone, and the loss is reported in the table. For professional and scientific services the loss is large enough that its numbers should be read with that in mind.
- Connecticut abolished its counties in 2022, and the replacement planning regions do not nest inside them and carry no BEA county GDP before 2024. So the window here stops at 2019, which also puts it before the pandemic deformed county output. The 2020–2023 window is reachable only with Connecticut dropped and is not mixed in.
- The MSA vintage in the 2007 paper is never stated. The metropolitan map was redrawn in 2003, 2013 and 2018. The delineation used here is the March 2020 one, which is certainly not the one that paper used, so the count of cities differs and a difference in count is a different sample rather than a discrepancy.
- The optimiser is a heuristic, as Openshaw said of his own. It finds a good local optimum, not a proved global one, so every range it reports is a lower bound on the true achievable range. Where restarts disagreed, the spread is in the table.
- The person-model's evidence measure does not transfer to money. Its exponent is invariant to whether output is counted in dollars or thousands of dollars, and is reported. Its information criterion treats each dollar as an independently assigned token, which is not a claim anyone should make about dollars, so that column is not used to adjudicate anything.
three results that did not come out as expected
Check it
Everything here is reproducible from a clean checkout. The lab is
research/urban-scaling-boundary/; the sources, including the verbatim transcriptions of
Bettencourt's Table 1 and Openshaw's Tables 11 and 12, are in its sources/ directory,
each with a note saying what could not be reached and why.
bash research/urban-scaling-boundary/fetch.sh # the public files, about 60 MB python3 research/urban-scaling-boundary/prep-delineation.py python3 research/urban-scaling-boundary/prep-commuting.py node research/urban-scaling-boundary/analyse.mjs # the cross node research/urban-scaling-boundary/zone-experiments.mjs # Iowa, and the zoned exponent node research/urban-scaling-boundary/profile-audit.mjs # is the optimiser telling the truth? node verify-a-million-or-so-exponents.mjs # every figure on this page
The three engine files this page runs in your browser
(engine/cluster.mjs, engine/fit.mjs, engine/zoning.mjs) are
byte-for-byte copies of the ones the analysis ran, and the verifier asserts it. The exponent you get
by dragging the slider is computed by the same code that produced every number in the tables.
Sources. Bettencourt, Lobo, Helbing, Kühnert and West, PNAS 104(17):7301–7306 (2007), PMC1852329. Arcaute, Hatna, Ferguson, Youn, Johansson and Batty, J. R. Soc. Interface 12:20140745 (2015), preprint arXiv:1301.1674. Leitão, Miotto, Gerlach and Altmann, R. Soc. Open Sci. 3:150649 (2016), preprint arXiv:1604.02872. Openshaw and Taylor, "A million or so correlation coefficients", in Wrigley (ed.), Statistical methods in the spatial sciences (Pion, 1979), pp. 127–144, quoted through Openshaw, The Modifiable Areal Unit Problem, CATMOG 38 (1983). Yule and Kendall, An Introduction to the Theory of Statistics (Griffin, 1950), p. 312, quoted through the same.
Data. BEA regional accounts CAGDP2, CAINC1, CAINC4, CAEMP25N. Census Bureau
population estimates, county gazetteer, county adjacency file, CBSA delineation (March 2020) and
commuting flows (ACS 2011–2015). County presidential returns 2020 via the MIT-licensed tonmcg
mirror. County outlines from us-atlas, built on Census cartographic boundary files. All public
domain or openly licensed; every URL is listed in the lab's fetch.sh.