One row, then the whole surface

Who You Decide to Count

The Census Bureau printed one earnings ratio. The same row also supports another. Then 7,776 disclosed analysis choices spread much farther than the survey's sampling margin.

Which number, from which release. Everything here is income year 2023, collected by the Current Population Survey ASEC of March 2024 and printed in Census report P60-282, Income in the United States: 2023, September 2024. The Census has since published P60-286 for income year 2024, whose Table A-6 prints 0.809 for 2024 and reprints 0.827 for 2023. The two releases agree about 2023. This page reproduces the 2023 figure and never uses the 2024 one.

Full-time, year-round

0.827

$55,240 for women divided by $66,790 for men.

All workers with earnings

0.748

$42,110 for women divided by $56,280 for men.

Same survey. Same year. Same table row. The first ratio is the published headline.

This page will not nominate a replacement headline number. It reproduces the published point, then shows the surface created by stated choices. The refusal is deliberate.

The anchor

First, make 0.827 again

The browser is loading every CPS ASEC record with nonzero 2023 earnings. It applies the Census rule: $2,500 earnings intervals, linear interpolation inside the median interval, and a $250,000 plug for the open top interval.

Male FTYR median

loading

paper: $66,790

Female FTYR median

loading

paper: $55,240

Recomputed ratio: loading

The medians recompute to the printed tens of dollars. The unrounded ratio is 0.82709, which does not equal 0.827 exactly; it rounds there. That residual stays visible below. Table A-6 of the same report prints a 90 percent margin of error of 0.0107 on the 0.827 ratio.

The operable surface

Move one choice

Each result is 100 × (1 − exp(βfemale)) from a log-earnings regression. Positive values mean lower earnings for women in that specification. This is not the Census median estimator, so the paper's 17.3-point gap is plotted as a separate reference, not disguised as an OLS cell.

Occupation and industry are disputed controls. They can compare workers in more similar jobs, but sorting into jobs may itself be part of what is being measured. A higher rung is not automatically a truer rung. Experience here is potential experience, not actual work history. The ladder is nested, so rung 4 and above always carry log weeks and log hours, and rung 7 and above always carry marital status and children. Two of these rungs switch another control off, and the instrument says so on screen when you reach them.

Your selected specification

loading

loading

Same specification, family restriction changed

loading

loading

Loading the disclosed surface.

Where the instrument goes quiet

Two controls that switch each other off

One. The hourly wage is annual earnings divided by weeks and then by hours, so in logarithms log(hourly) = log(annual) - log(weeks) - log(hours). From rung 4 onward the ladder puts log weeks and log hours on the right-hand side of the regression with free coefficients. Once they are there, moving those two terms from the left-hand side to the right changes nothing about the female coefficient. Choosing "hourly" and then controlling for hours is not a stricter comparison. It is the annual comparison, relabelled. Across the 4,320 cells at rung 4 and above, the three earnings measures differ by at most 0.0000000006 percentage points, which is arithmetic noise, not a result. Across the 3,456 cells at rungs 0 to 3, where hours are not controlled, the same switch moves the answer by between 1.21 and 14.25 points, so the axis is real everywhere else.

Two. The never-married, no-children sample has already applied the family restriction as a filter. Inside it, marital status and number of children take one value each, so the rung 7 control columns are constant and the fit drops them. Across all 432 restricted cells, adding that rung moves the estimate by at most 0.0000000002 percentage points. In the unrestricted sample the same rung moves it by between 0.0005 and 0.68 points. A reader who ticks the family controls here and sees nothing happen has not learned that family structure does not matter. They have learned that the sample already absorbed it, one control away from where they are looking.

Both settings stay reachable, because hiding them would hide the point. The instrument names them on screen the moment you select them.

Loading 7,776 specifications.

minimum
...
median
...
maximum
...
women ahead
...
under 2 points
...

116 of the 7,776 cells land under two percentage points, and 45 put women ahead. None of them is a finding on its own; each is one cell, and the readout above pairs every under-two selection with the same specification minus the family restriction.

Locating the published point.

The depth layer

What carries the spread?

On a complete balanced grid, one-way main effects partition cleanly from their interaction remainder. The bars are computed from all 7,776 cells, not from the current selection.

Choice spread ÷ printed 90% margin

...

The curve spans ... percentage points. Table A-6 prints a 0.0107 margin of error for the 0.827 ratio, equal to 1.07 percentage points on this scale.

On the shipped grid the family restriction carries 56.11 percent of the variance and the covariate ladder carries 20.67 percent, with 16.91 percent left in the interaction remainder. Adding education alone raises the anchor OLS gap from 17.76% to 23.57%, because women in this sample are better educated. The ladder is not a staircase toward one answer. On the anchor cell it runs 17.76, 23.57, 23.08, 22.42, 20.19, 17.64, 18.71, 18.42, 18.48. That is neither monotone nor convergent, it passes below where it started, and its largest single step is the one that makes the gap wider. Showing only the two endpoints would report a drift from 17.76 to 18.48 and conceal a 5.81-point widening in between. Leaving education off the ladder would not be the cautious choice; it is the choice that makes adjustment look smallest.

The same question asked of itself

Which axis wins is also a choice

That 56.11 against 20.67 is a fact about this grid, not about the world, and this page is not entitled to state it as a verdict without showing how far it travels. It does not travel far: delete one option, the never-married and childless samples, and the covariate ladder becomes the largest axis on the remaining 3,888 cells at 46.87 percent. Here is the same comparison recomputed on subsets of the page's own surface. The sample axis here is every combination of worker definition, age band and family restriction, twelve in all; the covariate axis is the nine ladder rungs; the figure is the mean spread of the female coefficient along one axis holding every other choice fixed, in log points.

gridsample axiscovariate ladderwinner
Recomputing from all 7,776 coefficients.

Recomputing.

Three asymmetries this page owes the reader, and none of them favour the answer above.

First, the participation margin is not on this grid and cannot be. Every one of the 7,776 cells is a regression on the logarithm of earnings, and the logarithm of zero does not exist, so every cell has already conditioned on having earnings at all. The universe here is the 73,413 people aged 15 or over with nonzero earnings out of 144,265 person records in the source file. Whether an adult has earnings in the first place is the largest thing the sample axis can do, it is heavily gendered, and no log-earnings grid, this one included, can put it on the board. Any figure this page prints for the sample axis is a floor.

Second, the two axes are not the same size. Twelve sample definitions against nine rungs is not a fair race, and the fair version has to be constructed. Matching the count exactly, by taking every choice of nine sample definitions out of the twelve and setting each against all nine rungs, the sample axis still wins: recomputing.

Third, the ladder is nested, so this page cannot run the cleanest test of all. The hours-and-weeks control enters at rung 4 and every rung above it inherits it, which means there is no version of this grid with the deep rungs but without that control. A grid that could drop it separately would probably shrink the ladder further, so the ladder's figure here is, if anything, generous to the ladder.

Checking which choices are genuinely inert.

The check: what the browser actually verified

Table A-7 reproduction

quantitypaperlive
male FTYR median$66,790...
female FTYR median$55,240...
ratio0.827...
male weighted count68,470k...
female weighted count52,850k...

Waiting for data.

Parse and grid checks

144,265 source rows → loading

trimmed weight sum: loading

specifications: loading

dead-control audit: loading

All 7,776 grid cells are defined. Structurally constant nuisance columns are removed by the rank-revealing fit. If a future data change empties a cell, the page prints: No estimate: this specification has no eligible rows. A solve that cannot identify the female contrast is refused, never converted into a number.

The screenshot trap, recomputed live

... for never-married, childless, age 25 to 34, FTYR, hourly, raw. ...

... for the same specification without the family restriction. ...

What the solver is, and what it is not

The deepest rung of this surface is itself exactly rank deficient. The seven civilian class-of-worker categories mark the same rows as the eleven major occupation categories, so their dummies sum to the same column and one direction of the design is unidentified. All 864 rung 8 cells carry it. That direction has exactly zero weight on the female dummy, so the female contrast is still identified, and a rank-revealing orthogonal factorization returns it. On the full-time year-round sample aged 15 and over, a Cholesky factorization of that same design's normal equations also succeeds and hands back a plausible number, which is the point: a factorization that does not complain is not a check. Shipped coefficients were reproduced from the shipped binary by two independent stable solvers, a float64 SVD and a column-pivoted Householder QR, to within 0.000001 log points, which is the whole-person weight rounding described below.

Choices this grid does not vary

Education-years mapping; cubic potential experience; children capped at three; race capped at five categories; Census division instead of state; OLS instead of an Oaxaca-Blinder decomposition; detailed occupation and industry; 160 replicate weights; Census disclosure swaps; the unobservable true upper tail; and every adult with no earnings at all. There is a second official United States headline for this quantity, the Bureau of Labor Statistics usual weekly earnings series, and it is not on this page because we could not reach the source to verify a published value.

Data scope, precomputation, and the numerical refusal

The shipped 2.13 MB binary contains all 73,413 people age 15 or older with nonzero earnings, including 104 negative earners needed to reproduce the Census medians and counts. Regressions require a logarithm, so their universe is the 73,309 positive earners. Weights in the shipped browser file are rounded to whole people; the precomputed grid used the source's two-decimal weights. This changes the paired raw gaps by less than 0.0003 percentage points and neither displayed hundredth.

The 7,776 coefficients were precomputed in float64 from these rows and loaded as a local data artifact. The browser recomputes the anchor and screenshot-trap pair directly from the binary, then recomputes every curve summary, percentile and variance share from all coefficients with the shared multiverse kit. This is narrower than a live 14,688-regression browser sweep. It is disclosed here because pretending otherwise would be worse than the reduction.

Detailed 522-code occupation and 263-code industry controls are not part of the default curve. Measured on this file, that design is 833 columns; on full-time year-round workers aged 25 to 54 it has 36,755 rows and rank 828 after structurally empty columns are dropped. Float64 normal equations on it return a female coefficient of +32.81 while a float64 SVD and a column-scaled LSMR agree on -0.1448. The normal-equation output is rejected as a numerical failure, not shown as a result, and the framing that single precision was the problem is withdrawn: precision is not what fails here, conditioning is.

The shipped design is not innocent either. Its deepest rung is exactly rank deficient by one, because the class-of-worker dummies and the major occupation dummies both mark the rows that carry an occupation code, and on full-time samples one class-of-worker category is empty as well. On the full-time year-round sample aged 15 and over the deepest design is 51,588 rows by 74 columns, 73 of them non-empty, of rank 72, and float64 reports its condition number as about 4e20. The unidentified direction places no weight on the female dummy, so the coefficient this page reports is unique; the verifier proves that dependency exactly from the shipped binary rather than inferring it from a condition number.

Sources and responsibility

Source: U.S. Census Bureau, Current Population Survey, 2024 Annual Social and Economic Supplement, covering income year 2023. Published anchor: Gloria Guzman and Melissa Kollar, Income in the United States: 2023, Current Population Reports P60-282, September 2024, Tables A-6 and A-7. Standalone Table A-7. The successor report, P60-286, prints 0.809 for income year 2024 and reprints 0.827 for 2023; nothing on this page comes from it. The data is a U.S. federal government work under 17 U.S.C. §105. Conclusions from this analysis are the responsibility of Artificial Wasteland, not the Census Bureau.