Experimental design · thermal physics

The Demonstration That Cannot Fail

Every science museum has the bench: a metal plate and a wooden one, at the same temperature, and the metal feels colder. It is a good demonstration, and it settles one question completely. It also cannot distinguish any of the four explanations people give for it, because the five materials on that bench are ranked in the same order by all four. This page computes what the demonstration can and cannot decide, searches the cupboard for specimens that would fix it, propagates the literature's own disagreement through the candidates, and finds that exactly one survives, worth about a sixth of a kelvin.

Put your hand on a steel table leg and then on the wooden top. The steel feels colder. It is not colder: a thermometer touched to each reads the same. Our own page on this, The Cold That Isn't There, walks through why, and the answer it gives is the accepted one. The sensation is not a temperature reading. Skin has no thermometer; what the cold receptors report is the skin's own surface temperature falling, and how fast. The number your nerves meet is the contact temperature, the value the interface jumps to the instant two bodies touch:

T_contact = (e_skin · T_skin + e_object · T_object) / (e_skin + e_object)

where e = √(k·ρ·c) is the thermal effusivity: conductivity, density and specific heat rolled into one number saying how greedily a surface takes heat. Steel's effusivity is roughly twelve times skin's, so the interface is dragged most of the way down to the steel's temperature. Oak's is a third of skin's, so skin barely moves. That is correct physics, and it is what the physics-education literature says too.

The question here is a different one, and more awkward: does the demonstration on the bench establish it?

Six accounts of cold, and a bench that agrees with all of them

Ask people why the metal feels colder and you get a small, stable set of answers. Written so that each predicts an ordering, which of two same-temperature objects should feel colder, they are:

the rival accounts

Only H_e predicts a number in kelvin. The other five predict a ranking and nothing more, and that asymmetry matters later. Now the specimens. The bench everybody uses is metal, more metal, glass, wood, foam:

the classic bench, computed from published property values

specimenkρρce = √(kρc)α = k/ρcfelt as

Conductivity k in W/(m·K), density ρ in kg/m³, volumetric heat capacity ρc in J/(m³·K), effusivity e in J/(m²·K·s½), diffusivity α in m²/s. "Felt as" is the contact temperature against skin at 33 °C with every specimen at 20 °C.

Read down the k column, then e, then ρc, then ρ. Same order, four times: copper, steel, glass, oak, foam. The Spearman rank correlation between any two of those four columns, across these five specimens, is exactly 1.000.

Whatever this bench does, it does identically under all four accounts. No result it could produce would favour one over another.

What the bench can and cannot decide

Below is the whole picture. Every pair of accounts gets a cell, green if some pair of specimens on the bench would make the two predict different outcomes, red if nothing on the bench can tell them apart. Add and remove specimens and watch it change.

specimens on the bench

0.00 K

account pairs separated

confounded

specimens on the bench

Start from the classic five and the picture is stark. Nine of the fifteen account pairs are separated, which is genuinely good, because the nine include the one that matters most. Everything on the bench sits at one temperature, so "cold is a property of the object" predicts no difference at all, and the bench delivers a spread of more than eleven kelvin. That misconception dies on contact, by a margin no instrument could miss. The demonstration earns its place in every museum for that alone.

The six red cells are every pair among conductivity, effusivity, heat capacity and density. Those are the four that people actually argue about. The demonstration meant to teach you it is effusivity cannot tell effusivity from the "metal is a good conductor" answer it was built to correct.

And it is not an artefact of which handbook we opened. Published property values disagree, sometimes wildly, so we propagated the full published range for every material through the comparison. Three of the six confounds survive that intact: conductivity against effusivity, conductivity against density, and effusivity against density are confounded on this bench for every value inside every published range. The other three escape through a single crack, copper against mild steel, whose volumetric heat capacities overlap once you admit the real spread in steel alloys.

What would break it, and why the cupboard is nearly empty

To separate conductivity from effusivity you need a specimen pair the two rank in opposite orders: material A with the higher conductivity but the lower effusivity. Since e = √(k·ρc), that requires A's volumetric heat capacity to be smaller than B's by more than enough to overturn its conductivity advantage.

That is hard to arrange, because ρc barely varies. Across nearly every solid it sits between about 0.02 and 3.9 million joules per cubic metre per kelvin, and once you exclude foams the range is about three. Conductivity, over the same materials, varies by a factor of ten thousand. The conductivity term almost always wins, so the two orderings almost always agree.

Almost. Here is every pair in a pool of twenty-six materials that inverts them, with the published ranges propagated through. A verdict of robust means the inversion holds for every value inside every published range; broken means some legitimate published choice destroys it; incomplete means it holds but at least one range was never established, so the verdict is optimistic and is not allowed to count.

every conductivity-against-effusivity inversion in the pool

the pairverdicteffusivity, Aeffusivity, Bmarginfelt gap

Ten candidates. Nine die. The one that dies most instructively is window glass against water, which is the repair anybody would reach for first, and which we believed for several hours. Water conducts heat worse than glass and stores about twice as much per unit volume, so on the standard European value for float glass, k = 1.0 W/(m·K), water's effusivity comes out higher and the inversion is real. But the standard heat-transfer textbook, Incropera, puts soda lime glass at k = 1.4, and at that value the inversion vanishes. The published range for one of the most common substances on earth is 0.94 to 1.46, and the answer to the experiment flips inside it. The discriminating experiment is undecided by the property data it depends on.

What survives is fused silica against water, and it survives for a reason worth stating: those are the two best-characterised substances in the cupboard. Water is pinned to a tenth of a per cent in density and under one per cent in conductivity by an international formulation; fused silica is a laboratory material with a narrow published spread. The discriminating pair is not the pair that is most different. It is the pair that is best known.

And then the finger arrives

The surviving inversion is worth at the skin. Is that enough?

There is a direct answer in the psychophysics literature, and it is brutal. Jones and Berris put people in a two-alternative forced choice on turned rods of copper, brass, nickel, stainless steel and nylon, with texture cues removed. Subjects identified nylon at 92 to 96 per cent. They could not discriminate any metal pair. Copper against stainless steel came out at 42 per cent correct, which is chance. A clinical study from 1974 using copper, stainless steel, glass and PVC disks had already found the same thing.

the anchor

effusivity ratio, copper to stainless

contact-temperature gap

correct, in the published test

The effusivity ratio and the gap are computed here from the property table; the percentage is the published result. A pair five times apart in effusivity, and more than a kelvin apart in contact temperature, is invisible to a fingertip.

So the fingertip's discrimination threshold for this task is worse than , and the one repair that survived scrutiny is worth : smaller by a factor of . Drag the resolution slider up in the matrix above to that value and the conductivity-against-effusivity cell goes red and stays red. It is not that the bench needs better specimens. The largest conductivity-against-effusivity inversion anywhere in the pool, robust or not, is still below the threshold. No collection of these materials, of any size, can make this demonstration discriminate by hand.

It can be made to discriminate by instrument. A fine thermocouple or a decent thermal camera resolves a sixth of a kelvin without complaint. That is the honest conclusion, and it is not a small one: the deconfounded version of this demonstration is out of reach of the sense organ the demonstration is performed with. It is further out of reach than that, in fact, for a reason the next section had to go and find.

The smallest bench that could come out wrong

Given a cupboard and a set of accounts you want to tell apart, what is the smallest collection of specimens that does it? Toggle the accounts and the page solves it exactly, by trying every subset in increasing order of size.

accounts in play

specimens, exact minimum

the one-at-a-time rule

On the cupboard a demonstrator could actually assemble, the classic five plus the two candidate repairs, the answer with all six accounts in play is four specimens, and the surprise is which four: window glass, foam, water and fused silica. There is no metal on it. Copper and mild steel, the specimens the whole demonstration is built around, turn out to carry the least information: they are far apart from everything, but they are far apart in the same direction under every account, which is exactly what a specimen must not be.

The natural way to build a bench is to add whichever single specimen tells you the most. Run that rule and it spends five specimens where four suffice, and the one it wastes is copper, chosen first because a metal looks informative on its own. A specimen is worth nothing alone; it earns its keep only in company, and a rule that scores specimens one at a time is structurally poor at seeing that. Scoring specimen pairs instead recovers the optimum here.

This is a known problem, and naming it correctly matters. When each hypothesis assigns each specimen a binary outcome, choosing the smallest distinguishing set is MINIMUM TEST SET, catalogued as problem SP6 in Garey and Johnson, NP-complete, and reducible to set cover, which is where the greedy algorithm's logarithmic guarantee comes from. But the accounts here predict orderings, so a single specimen carries no observable at all; the atom of observation is a pair of specimens while you pay per specimen. That is a different problem, SET COVER WITH PAIRS, and the standard greedy guarantee does not transfer to it. The exact answers above are found by enumeration, not by greedy, for that reason.

The finger is a disc, not a plane

Before going further there is a correction to make, and it costs us most of what we had just found.

The contact-temperature formula, and everything computed from it above, is the solution for two half-spaces: infinite flat bodies meeting over an infinite plane. A finger is not that. It is a patch about eight millimetres across, and outside that patch the specimen's surface is touching nothing. So heat runs sideways inside the specimen, into material the finger never touches, and the finger gets colder than the half-space formula says.

How much colder depends on how far heat can run, which is to say on the specimen's conductivity. So the correction is not a constant offset that cancels out of a comparison. It is largest for exactly the materials that conduct best.

an infinite plane against a real fingertip, at one second

specimenhalf-space8 mm fingerdifference
solving the axisymmetric heat equation, a moment

Area-averaged interface temperature under the contact disc, from an axisymmetric finite-volume solve in (r, z). Its validation, including the block limit, the one-dimensional limit, grid convergence and an energy audit, is in spread.test.mjs.

Water is barely touched by this, at , because water conducts too poorly to reach anywhere. Fused silica loses , because it conducts twice as well. And those two are the discriminating pair, so the correction eats the thing that separated them:

the one surviving repair, corrected for the shape of a finger

gap, infinite plane

gap, 8 mm finger

times overstated

So the final number is not a sixth of a kelvin. It is , against a fingertip that has been measured to be at chance on . The idealised model that everyone, ourselves included, uses to argue that effusivity is the right account also overstates the size of the only experiment that could test it, by a factor of . That the measured skin response in these experiments comes out smaller than semi-infinite theory predicts is a known and unexplained complaint in the psychophysics literature; some of it is this.

The axis nobody uses

All six accounts share something easy to miss: each says the sensation is a property of the material. Copper is a cold material; wood is a warm one. Every one of them assigns a number to a substance and stops.

So here is a second experiment. It needs no rare specimen and no fine instrument. Take one material, make two specimens of it, one thick and one thin, rest the thin one on something insulating, and touch both.

contact temperature against time, under an 8 mm finger

the slab, solved in two dimensions the same slab, one-dimensional model a block of the same material a block of oak

A block of aluminium holds its contact temperature near indefinitely: a fingertip cannot dent it. A sheet of ordinary kitchen aluminium foil, sixteen microns thick, resting on foam, reads after a second. It is warmer than a block of oak. Same metal. An eleven-kelvin swing, wider than the whole classic bench manages with five different materials.

Nothing in any of the six accounts permits that, because none of them contains a length. The entire family is refuted at once, with one material, a pair of scissors and a kitchen drawer.

The mechanism is not the one the one-dimensional picture suggests. In one dimension a thin slab simply runs out of cold, so it should be limited by the heat it holds, ρcL, and stainless steel should overtake aluminium once both are thin. It does, in that model. Under a real finger it does not: aluminium stays colder and the gap widens, because a thin sheet of a good conductor is not a small reservoir, it is a heat spreader, reaching sideways into metal your finger never touches. Stainless steel cannot do that.

Which produces the sharpest thing on this page. In one dimension, every metal foil is within a quarter of a kelvin of every other: they are all exhausted, and the model says they all feel the same. Under a real finger they spread over nearly a kelvin and a half, and copper foil, the best conductor, is the coldest. In the foil regime, conductivity really does decide. The account the demonstration was built to correct is not wrong; it is right about a case the demonstration never shows you, while the case it does show you cannot tell the two accounts apart.

What this generalises to

The specific result is small: a physics demo has a confound. The shape of it is not.

A demonstration is evidence only if some outcome it could have produced would have counted against the thing it claims. That is Popper's requirement in the form most people met it, and it has been sharpened a great deal since: it is not enough that a claim be falsifiable in principle, the particular test you ran must have had a real chance of coming out the other way. Deborah Mayo's name for the property is severity, and her name for its absence is better still. A result that agrees with a claim, produced by a method practically incapable of finding flaws in that claim even if they exist, is bad evidence, no test. The metal-and-wood bench is a severe test of one account and a zero-severity test of four others, and from the outside the two look identical. Nobody is lying. The demonstration works, the audience learns something true, and the thing they were told it proved is not the thing it proved.

What makes this tractable rather than merely worrying is that the property is computable in advance. You do not need good instincts about confounds. Write down what each rival account predicts for each candidate specimen, and "could this bench have come out otherwise" becomes arithmetic: rank the specimens under each account and look for an inversion. A Spearman coefficient of exactly 1 between your explanation and its rival is the compact statement of a confound, and you can compute it before you build anything. The harder half, which this page kept running into, is that the arithmetic is only as good as the property values, and propagating their real disagreement killed nine of our ten repairs.

We built this because it caught us. The confound is in our page. The Cold That Isn't There names its target precisely, says the answer is effusivity and not conductivity, and then demonstrates it on five specimens that cannot tell those two apart. Its own cross-check table already contained water, computed and correct, one row away from the tiles and never put on the bench.


What is new here, and what is not

That the correct account is effusivity rather than conductivity is published and settled in the physics-education literature; Marín argued it in The Physics Teacher in 2006 and Oss developed it in the European Journal of Physics in 2022. Neither proposes a specimen that would discriminate the two accounts, and neither notes that the usual specimen sets cannot.

We searched for prior statements of the co-ranking observation and did not find one, in physics education, in museum-exhibit literature, or in the psychophysics of thermal perception. We also found no study that manipulates conductivity and effusivity orthogonally: every real-material experiment uses natural materials in which they covary, and every thermal-display experiment renders a simulated contact-temperature trajectory, which presupposes the effusivity model rather than testing it. One near miss deserves recording: Ho's 2017 review contains a table in which foam and acrylic have essentially equal conductivity and effusivities differing by a factor of 2.7, an inversion sitting in print, unremarked.

A clean negative worth keeping: the psychophysics literature does not agree with itself about which property people perceive. Jones and Berris concluded heat capacity; Bergmann Tiest and Kappers titled their paper diffusivity while controlling heat-extraction rate; Ho's review says contact coefficient, which is effusivity. The physics-education literature's confident answer is better supported theoretically than experimentally, which is itself part of the story of this page.

Limits, and what we could not establish

The check

Everything above is recomputed in your browser as you read, by the same two files the offline verifier loads. Nothing on this page is typed in. These assertions run now, on your machine, on the numbers actually being displayed:

live self-check

The offline verifier runs a superset of these, including the negative controls and the solver validation that cannot run in a page:

node research/the-demonstration-that-cannot-fail/verify.mjs

Sources

Property values, with the source for each specimen shown in the table below, and then the literature the argument rests on.

where every property value came from