Nobody Compared Three
In 2014 a paper in Science put the number of smells a human can tell apart at more than a trillion, and the figure went everywhere. Its trial record is public. This page rebuilds the arithmetic, which reproduces all six published numbers exactly, and then asks a different question of the same data: how large a set of odours did the experiment actually show to be distinguishable from one another?
The answer is two. Not two thousand, not two million. Two. And that is not a criticism of the noses, which did fine. It is a fact about the shape of the experiment, and it can be read straight off the trial record without any model at all.
What kind of number a trillion is
The paper opens by placing smell alongside the other senses. Researchers
estimated that humans can distinguish between 2.3 million and 7.5 million colors
and about
340,000 tones, it says, and nobody knows the olfactory figure. Those comparison numbers are
counts of a specific object: a set of stimuli every one of which can be told apart from every
other one. Take any two colours out of the millions and they look different. That is what makes
the number a number.
To measure smell the authors built mixtures. From a palette of 128 intensity-matched odorant molecules they made mixtures of exactly 10, 20 or 30 components, and they paired mixtures that shared a controlled fraction of their ingredients. Twenty-six subjects each ran 264 triangle tests: three vials, two holding the same mixture, one holding a different one, pick the odd one out. The more two mixtures overlapped, the harder that got. Then the authors fitted a straight line through the results, read off the overlap at which half the subjects could still tell a pair apart, and turned that number into a count of mixtures by packing spheres.
Everything in that chain is public: the formulas are in the supplementary materials, and the trial-by-trial record, every vial's ingredients and every subject's answer, is in Table S2. So the whole thing can be taken apart.
The machine, rebuilt
Three formulas do the work. Around any mixture X, the set of mixtures that differ from it in exactly R components is sphere(R) = C(N,R) · C(C−N,R): choose which R of X's components to swap out, and choose their replacements from the C−N components not in X. Summing that from 0 to R gives a ball. Divide the total number of mixtures by the size of one ball and you get how many balls fit, which the paper takes as the number of mixtures that can be told apart:
disc(D) = C(C,N) / ball(R), D = 2R = N − O
The measured resolution enters as O, the overlap in components at which discrimination fails. The fitted line y = −0.81x + 91.45 crosses 50% of subjects at 51.17% overlap, and the second line y = −0.77x + 94.22 crosses at 57.43%. Those two percentages are the only empirical input the packing calculation ever receives.
One wrinkle: 51.17% of 30 components is 15.351, so the radius R is 7.3245, and a ball of
fractional radius does not exist. The paper's supplement says the values were
extrapolated from the graphs using Mathematica
, and those graphs are semi-logarithmic, so
reading a point off the line between two whole-number radii is log-linear interpolation. Do that,
and the machine gives back every number the paper printed:
Reproduction: formulas (1) to (3), against the published figures
Six for six. That is worth stating plainly because it means the object below is the real estimator and not a paraphrase of it, and because the exact provenance of the trillion has not, as far as I can find, been pinned down in print before: it is formula (3) at C = 128, N = 30, with the radius interpolated between 7 and 8.
Turn the two knobs it was never a function of
C is the size of the odorant palette and N is the number of components per mixture. Neither is a property of a nose. Both were chosen by the people running the study, and the answer depends on them steeply. Hold the psychophysics exactly where the subjects left it, at 51.17% overlap, and move only the design:
The curve is disc(D) against overlap for the current palette and mixture size, on a log axis from 1 to every possible mixture. The dot is where you are standing.
This dependence is not news, and it is not mine. Gerkin and Castro published it in 2015, noting
that the estimate grows as roughly the thirtieth power of the palette size, so that a library the
size of a commercial flavour catalogue, about 2,000 chemicals, would put the figure near
1041, implying a unique olfactory percept for each carbon atom on earth
. Their
point stands entirely and the slider only lets you feel it move. Two small things about the
illustration, though, now that it can be recomputed. The exponent is not thirty: the numerator
C(C,N) does grow as the thirtieth power, but the ball in the denominator grows with the palette
too, and over that range the ratio comes out at 23.2. And
formula (3) at the paper's own threshold gives
at 2,000 chemicals, a little over an order of magnitude below the figure they quote. Neither
changes anything they were arguing.
The authors, for their part, printed three of these numbers themselves, one for each mixture size, and offered the mixture-size dependence as a reason their figure was conservative: real smells can have more than 30 components, so the true count must be higher. That is a fair argument about the world. It is not an argument that the arithmetic is measuring the nose.
The board
Now set the arithmetic aside and look at what happened in the room. Thirteen kinds of mixture pair, twenty pairs of each, 260 tests. Here they all are. Each mark is one test: two dots for the two mixtures, a bar joining them, and the bar is bright where 14 or more of the 26 subjects found the odd vial, which is the paper's own criterion for a pair being discriminable.
Thirteen rows, one per stimulus type, labelled by mixture size and per cent overlap. Twenty tests per row. Bright bars are the 148 pairs discriminated significantly above chance; dim bars are the rest. The single thin arc joins the only two tests in the whole study that share a mixture.
The picture is the argument. Nothing touches anything. Two mixtures per test, 260 tests, 520 mixture slots, and 519 distinct mixtures: every mixture in the study was used in one test and one test only, with a single exception. A ten-component mixture of eucalyptol, eugenol, menthol and seven others appears in test 161 at zero overlap, where 20 of 26 subjects placed it, and again in test 80 at 30% overlap, where 18 of 26 did.
So the study contains exactly one chain of two links, and no closed triangle anywhere. Gerkin and
Castro noticed the pattern in 2015, writing that each mixture of a tested pair is used only
once in (Bushdid et al., 2014), in that pair alone, and never in any other pairs
, and using
it to argue that no map of perceptual space could be built from the data. That is very nearly
right, and the exception is worth having: one mixture was reused, and even it yields a path
rather than a triangle.
Two
2
Put a vertex on every mixture and an edge on every pair the experiment showed to be discriminable, and the largest set of odours demonstrated to be mutually distinguishable is the largest clique in that graph. Because no vertex has more than two neighbours and the only vertex with two lies on a path, the clique number is 2. It is 2 whether you use all 260 tests or only the 148 significant ones. Nobody compared three.
This is a fact about the design, not a failure of it. Certifying a mutually discriminable set of size k requires the k(k−1)/2 comparisons among its members, and 260 tests spent perfectly on that goal would have certified a set of at most 23; the 148 significant results, at most 17. Twenty-three is not an interesting number of smells. Spending the tests instead on a spread of overlap levels, which is what the authors did, is the right way to estimate a psychometric function. It simply leaves the count itself entirely to the model.
How long the direct check would take
It is worth seeing why the model is unavoidable rather than merely convenient. The subjects ran 264 tests over three visits with a median duration of 1 hour 6 minutes, which is 45 seconds per test. To certify a trillion odours mutually discriminable you would need triangle tests. At the study's own pace that is
So there was never a version of this experiment that measured the answer. The number had to come from a model, and the only question worth arguing about is which one. That is exactly where the dispute went.
Six numbers on one axis
Here is the whole quarrel on a single log scale, with each value coloured by what kind of thing it is: green for what the noses demonstrated, amber for what the arithmetic asserts.
A logarithmic axis from 1 to 1030. Marked, left to right: the set the experiment demonstrated; the set the paper's own criterion asserts among the study's own 30-component vials; the Gilbert-Varshamov floor; the paper's published figure; the rigorous Hamming ceiling at a whole-number radius; and the total number of mixtures of 30 from 128.
Two of those markers are Gerkin and Castro's. They identified formula (3) as the Hamming bound of
coding theory, which is an upper bound on how many codewords fit, and pointed out that the same
expression summed to D rather than D/2 is the Gilbert-Varshamov bound, which is a lower one.
Their Figure 6C states the consequence: using the paper's own resolution, the number
may be as small as ~10,000, and is guaranteed to be no larger than ~1 trillion
. Recomputing
that floor here gives , which is their ~10,000.
The direction matters because of what the paper says on its cover. The title is
Humans Can Discriminate More than 1 Trillion Olfactory Stimuli
and the abstract says
at least
. The supplement, describing the same formula, says:
This is an upper bound as it fails to count the “dead space” in the corners between spheres. Bushdid, Magnasco, Vosshall & Keller 2014, supplementary materials, Calculation of the Number of Discriminable Mixtures
Both sentences are in the same paper.
The correction you could not see
The authors did address this, and the story of how is worth telling. On 18 August 2016 Science posted a correction listing four assumptions the model had made without saying so. The third one reads:
Third, our model used a simplified approximation to obtain the total number of spheres that can be packed in a space by dividing the overall volume of the space by the volume of a single sphere. The actual calculation of the number of packable spheres using the “spherical code” problem establishes a rigorous lower bound a factor of 10 smaller than our estimate. Correction posted 18 August 2016, as reproduced verbatim by the corresponding author on PubPeer, February 2024
That correction was hard to see for eight years. In the corresponding author's own words on
PubPeer, the downloaded PDF was not revised to include the correction and the correction is
only visible on the online version of the paper
. The copy hosted by the authors' own
university, which supplied every sentence of the paper quoted on this page, carries no notice of
it. In February 2024, at the authors' request, Science reindexed the correction as a
formal Erratum whose entire new content is the announcement that it is now formal, and that
Erratum is dated eight days after the corresponding author posted the 2016 text on PubPeer.
Two things about that third assumption. First, it concedes the direction of the bound, which is the substantive point. Second, the number does not obviously survive contact with the standard result. A factor of 10 below 1.72 × 1012 is 1.7 × 1011. The constant-weight Gilbert-Varshamov bound at the same parameters, which is the rigorous lower bound for exactly this problem, is : not one order of magnitude below the estimate but about eight. The correction gives no formula and cites no source for the factor of 10, and the preprint it cites for its other claim contains no occurrence of “spherical code”, “packing”, “Hamming”, “Gilbert” or “lower bound”, and no equations at all. I could not find the calculation anywhere. That does not mean it does not exist, and I would genuinely like to see it; it means that as of this writing the repaired bound is the one number in this story with no published derivation.
The experiment nobody has run
Auditing is cheap. Here is something more useful: the smallest claim the model makes that could actually be checked, priced.
Among the 30-component mixtures this study physically mixed and put in front of people, take the largest set whose members all overlap one another by less than the paper's own 51.17% threshold. The answer is exactly 120 of the 160, and it has a proof rather than a search behind it: of the pairs among those vials, only 40 fail the threshold, those 40 are exactly the high-overlap pairs deliberately built inside single tests, and no vial appears in two of them. Delete one member of each and you have 120; you cannot do better, because each of the 40 forces a deletion.
The model therefore predicts that those 120 real vials, which already exist as recipes in Table S2, are mutually distinguishable. The median pair among them shares 7 of 30 components, where the paper's own fitted line predicts % of subjects can tell them apart. Checking the prediction takes 7,140 triangle tests, which at 45 seconds each is hours of one nose, or about times the load each of the original 26 subjects already carried. That is a real experiment, on the small side for a psychophysics lab, and it would move the demonstrated number from 2 to 120 or else fail informatively. Twelve years after the study, I can find no report of anyone running it.
The microbe, run again
The sharpest objection to the whole framework was Markus Meister's in 2015, and it is worth restating in his terms because it is the same shape as the graph above. The measurement licenses only that neighbouring stimuli differ. It says nothing about distant ones, so the same percept may recur across the space, and:
To obtain the number of discriminable odors, we need to determine the largest set of stimuli such that every stimulus can be discriminated from every other one, not just from nearest neighbors. Meister, On the dimensionality of odor space, eLife 2015;4:e07865
To show what that does to the estimator he built a positive control: a creature with exactly three responses. Pure odorants are attractants or repellents worth +1 and −1, a mixture's sum below −2 is “yuck”, above +2 is “yum”, in between is “meh”, and two mixtures are discriminable when the creature answers differently. Run the study's protocol on it and the estimator returns roughly 1012, while the truth, by construction, is 3. Press the button and watch it happen:
Not yet run.
Fraction of mixture pairs the creature answers differently, against how many of the 30 components the pair shares. The dashed line is the 50% criterion. Computed live in your browser from the same module the verifier uses, with the valences redrawn every repeat.
That redrawing matters, and getting it wrong cost me an hour. Hold one random assignment of
attractants and repellents fixed and the answer depends on how lopsided that particular draw
happened to be: a palette with 70 attractants and 58 repellents pushes every mixture's sum away
from zero, shrinks the “meh” band, and moves the crossing from 15 shared components
to 17. Averaged over repeats, as Meister does, the curve tops out at
against a closed form of
, which is his almost 2/3
, and crosses one half at
15 shared components of 30, which is his figure exactly. Feed 15 into formula (3) and it returns
, against the ~9·1011
he
reports. The replication holds.
The original authors replied to this, and their reply also contains something checkable. They
argued that the creature is a rigged test because it misbehaves at other mixture sizes:
The discrimination curve for 10 components never crosses 50%, and the discrimination curve for
60 components regresses back, so the threshold is non-monotonic.
That is a claim about a
model, so it can be settled by running the model, and they are right:
The toy microbe at three mixture sizes
At ten components the creature's curve peaks at barely above one half and falls below it as soon as a single component is shared, so there is no usable limen at all. At sixty the crossing comes back down in fractional terms. The threshold really is non-monotonic in mixture size, and that is a genuine defect in the toy. Note also a smaller thing: Meister puts the crossing at 30 components at 15 shared, and the reply's figure caption says 14. Neither is in error. The curve sits within a thousandth of one half at 15 shared components, so which side of the line it lands on depends on the random seed, and different seeds here give both answers.
But the defect in the toy does not repair the estimator, and this is the point the whole page has been walking toward. Meister needed a model to show that nearest-neighbour discriminability does not chain into a mutually discriminable set. The board above shows the same thing with no model at all: in the actual trial record, discriminability was never once measured in a way that could chain. There is nothing there to call rigged.
What is fair to conclude
Not that human smell is poor. Nothing here bears on that, and a 2017 review in the same journal argues at length that the belief in a feeble human nose is a nineteenth-century inheritance rather than a finding. Not that the trillion is refuted, either: the true count could be a trillion, or far more, or far less. The authors' own defence, that any exponential is sensitive to its inputs and that their aim was an order of magnitude rather than a figure, is reasonable as far as it goes, and their argument that olfactory space is high-dimensional may well be right.
What the trial record shows is narrower and harder to argue with. The experiment demonstrated
two. Everything above two is the model talking, the model's answer moves by dozens of orders of
magnitude when you move numbers the noses had no part in choosing, and the formula that produced
the headline is an upper bound wearing the words “at least”. A decade on, the field's
own flagship review still puts it the way the second critique's title did:
the exact number of distinct odorous molecules detectable and differentiable by the human nose
remains unknown
.
The check
- Recomputed from the primary sources, not quoted. Formulas (1) to (3) are implemented from the supplement in exact integer arithmetic and reproduce all six published figures to three significant digits. The trial record is parsed from Table S2 and reproduces the paper's own two counts, 227 pairs above chance and 148 significant, exactly.
- The structural claim. 519 distinct mixtures across 520 slots; maximum degree 2; maximum clique 2, by exact branch and bound over the whole graph, on all 260 tests and on the 148 significant ones separately.
- The 120-vial set is verified two ways: by the same exact clique search, and by the matching argument, which is checked rather than asserted (the 40 failing pairs are confirmed vertex-disjoint).
- Meister's control is re-implemented from his published description and its two reported figures are recovered. The reply's counter-claim is tested at three mixture sizes and confirmed.
- Assumed, not shown. That a mixture is identified by its set of components, which is the study's own convention (the three vials of a test are deliberately at different dilutions so intensity cannot be the cue). That 45 seconds per test, derived from the supplement's median visit duration, extrapolates to much longer runs; treat the sniffing-time figures as orders of magnitude.
- Could not obtain. The body text of the 2024 Erratum on science.org, which is
paywalled; the 2016 correction is quoted here from the corresponding author's own verbatim
posting of it, and its three-item reference list matches the Erratum's exactly. The
calculation behind the correction's
rigorous lower bound a factor of 10 smaller
: not in the preprint it cites, not found anywhere else. If it exists, this page is wrong about that one paragraph and I would like to know. - Not claimed as new. The knob dependence, the Hamming-versus-Gilbert-Varshamov correction and the 104 to 1012 bracket are Gerkin and Castro's. The positive-control method and the toy creature are Meister's. The no-reuse observation is Gerkin and Castro's; the clique number, the certification ceilings and cost, and the 120-vial experiment are this page's.
Verifier: verify-nobody-compared-three.mjs at the repository root. Data and analysis: research/trillion-odours/. The estimator and the microbe are the same module files this page loads.
Sources
- C. Bushdid, M. O. Magnasco, L. B. Vosshall, A. Keller, Humans Can Discriminate More than 1 Trillion Olfactory Stimuli, Science 343, 1370 (2014). doi:10.1126/science.1249168. Report and supplementary materials read from the author-posted PDF at rockefeller.edu; Tables S1 and S2 as redistributed in the repository below.
- M. Meister, On the dimensionality of odor space, eLife 4, e07865 (2015). elifesciences.org/articles/07865.
- R. C. Gerkin, J. B. Castro, The number of olfactory stimuli that humans can discriminate is still unknown, eLife 4, e08127 (2015). elifesciences.org/articles/08127. Code and the digitised Bushdid tables: github.com/rgerkin/trillion.
- M. O. Magnasco, A. Keller, L. B. Vosshall, On the dimensionality of olfactory space, bioRxiv 022103 (posted 6 July 2015). doi:10.1101/022103. Never published in a journal.
- Erratum for the Report “Humans can discriminate more than 1 trillion olfactory stimuli”, Science 383, eado6457 (2024). doi:10.1126/science.ado6457. The 2016 correction it formalises, quoted above from the corresponding author's posting at PubPeer.
- G. N. Dikeçligil, J. A. Gottfried, What Does the Human Olfactory System Do, and How Does It Do It?, Annual Review of Psychology 75, 155 (2024). doi:10.1146/annurev-psych-042023-101155.
- J. P. McGann, Poor human olfaction is a 19th-century myth, Science 356, eaam7263 (2017). doi:10.1126/science.aam7263. Cited here only for the sentence about human olfactory ability; nothing on this page depends on it.