Nobody Chose This Volume

You are in a full restaurant, leaning across the table, asking someone to repeat themselves for the third time. Nobody in the room is shouting on purpose. Every single person is raising their voice by exactly the amount that seems reasonable, and the room has settled at a level that no one in it would have picked.

Two facts, both measured, sit awkwardly together.

The first: put a person in noise and they get louder. Étienne Lombard reported it in 1911, and it is not a decision, it is a reflex. You do it while insisting you are not doing it. The rate is about half a decibel of extra voice for every decibel of noise.

The second: sound adds in energy, not in loudness. Two people talking make about three decibels more noise than one, not twice as much. Double the crowd, add three decibels. That is the arithmetic of incoherent sources and it does not care how you feel about it.

Hold those two together and the second one turns out to be wrong, in a room with people in it. Doubling the crowd does not add three decibels. It adds six. The extra three come from everybody answering the first three, and then answering each other's answers, all the way down. The room is an amplifier, and the gain is exactly two.

The loop, drawn

Here is the whole mechanism as two straight lines.

The rising line is what people do: the level you speak at, given the noise you are sitting in. It starts at 55 decibels when the room is at 45, which is where the reflex switches on, and it climbs with slope c, the Lombard slope. The other line is what the room does: the noise the room holds, given the level everyone is speaking at. Its slope is exactly 1, because a decibel in is a decibel out.

The room settles where the lines cross. Drag the slope.

Instrument one · the fixed point

Loop gain 1/(1-c)
2.00
Room settles at
71.0 dB
Everyone speaks at
68.0 dB
Per doubling of crowd
6.02 dB

Two things are worth staring at.

At c = 0 the lines cross low and the loop gain is 1: people would simply talk, and the room would be whatever the arithmetic of adding voices says it is. Every decibel of gain above that is the crowd answering itself.

Push c toward 1 and the lines become parallel. There is no crossing point. That is the runaway that people mean when they say a party gets out of hand, and the reason it does not actually happen is the whole of the reassurance available here: the measured slope is about a half, so the lines do cross, so the room does settle. It settles somewhere unpleasant, but it settles.

The 1959 paper that started this, W. R. MacLean's On the Acoustics of Cocktail Parties, concluded that there is a critical number of guests past which a party abruptly becomes loud. The measurements since do not show a cliff. They show a smooth ramp of about six decibels per doubling, all the way up, which is worse in an ordinary way: there is no threshold to stay under.

A room, and what it does to you

Now put numbers on it. Below is J. H. Rindel's prediction model from 2010, which is the above two lines solved for their crossing point and nothing more:

LN = 1/(1 − c) × [ 69 − 45c − 10 log10(A / NS) ]  dB

where A is the room's absorption area in square metres and NS is the number of people actually talking at any given moment, which is fewer than the number of people present, because at a table of four only one person talks at a time. Rindel calls the ratio the group size. At c = 0.5 the whole thing collapses to something you can almost do in your head:

LN = 93 − 20 log10(A / NS)  dB

Build a room. The absorption comes from the Sabine equation, so the two things you set are the volume and how long a clap takes to die away in the empty room.

Instrument two · build the room, then listen to it

Absorption area A
74 m2
Room noise
77.2 dB
You must speak at
71.1 dB
Vocal effort
raised
Signal to noise, 1 m
-6.1 dB
Conversation
insufficient
Acoustic capacity
19
Over capacity by
2.4x

Synthesised, not recorded: speech-shaped noise with syllable-rate envelopes, one voice per talker up to twelve and a steady bed for the rest. Your volume knob sets the absolute level, so treat the loudness as arbitrary. What is not arbitrary is the ratio: the near voice and the room are mixed at exactly the signal-to-noise ratio the model computes above. Turning the reflex off recomputes the room at c = 0, which is the same crowd not answering itself.

Two of those readouts deserve a word.

Signal to noise is the level of the person you came with, measured a metre away, minus the level of the room. Below about −3 decibels, the acoustics literature stops calling conversation sufficient. Note where that lands relative to the rooms you actually eat in.

Acoustic capacity is Rindel's proposal that a room should be labelled with the largest number of people at which talking still works, the way it is labelled with a fire limit. The rule of thumb is memorable:

Nmax ≈ V / (20 T)

Volume in cubic metres, reverberation time in seconds. It is a simplification, and it is worth knowing what of. Solve the signal-to-noise condition exactly, with the diners' own bodies counted as absorbers, and the divisor comes out at 20.29, not a round 20, for the group size and per-person absorption that paper assumes. Change the assumption about how talkative the room is and the divisor moves a long way: . The formula is honest; it is the group size that is soft.

The doubling that pays twice

Here is the part that is genuinely useful, and it falls out of the same factor of two.

Take an ordinary noise source, a ventilation fan say, and double a room's absorption. The fan pushes out the same power as before, the room holds less of it, and the level drops by three decibels. That is the textbook result and everyone in building acoustics knows it.

Now do it in a room full of people. The level drops by three decibels, so everyone stops straining, so everyone drops their voice by half of that, so the level drops again, and the whole series sums to six.

Instrument three · the same treatment, two rooms

A fan quietens by
3.0 dB
A crowd quietens by
6.0 dB
Ratio
2.00x

The same number, 1/(1 − c), governs the disease and the cure. It is why a crowded room is worse than you would predict, and it is why hanging absorption in one is better than you would predict. Acoustic treatment works on people better than it works on machines, because machines do not have opinions about whether they are being heard.

Ten real rooms

Hodgson, Steiniger and Razavi measured ten eating places in Vancouver in 2007, and Rindel tabulated them against his model. The table below is not copied from the paper. It is computed in your browser from the volume, the empty reverberation time and the number of seats, using the two equations above and nothing else, and then compared with what was actually measured with a sound level meter in the room.

Instrument four · the model against the meter

RoomV, m3T, sSeatsA, m2 PredictedMeasuredMissCapacityConversation

C: cafeteria. B: bistro. R: restaurant. S: senior residence dining room. Group size 4 for the first eight, 8 for the last two, and 0.5 m2 of absorption per person, exactly as the paper assumes.

Two things fall out of that table that are worth saying plainly.

The first is that the model works. Eight of the ten land within about a decibel, from a calculation with three inputs, against a real meter in a real room. Predicting the noise of a crowd from the shape of the room is not supposed to be that easy.

The second is the capacity column. Filled to their own seat count, eight of these ten rooms are past the point where the literature calls conversation sufficient. The two that are not are the dining rooms in the senior residences, and they pass for a reason that is not architectural: the paper needed a group size of eight to fit them, against three or four everywhere else. Fewer people talking at once. It is the quietest room in the study and the loneliest number in the table.

While reproducing this table, one small thing came up. The paper's discussion says the model lands within a decibel in eight of the ten cases. Recomputed at full precision, six do; two more sit at 1.06 and 1.10 decibels, and Rindel's own table prints both as “1.1”. The claim survives as eight of ten within 1.11 decibels. The model is fine. The sentence rounds. It is a small thing, but a check that never catches anything is not a check, and this one is reported by the verifier rather than smoothed over.

Where the model stops being true

The equation above contains no distances. Not one. It treats the room as a single tank of sound that every talker pours into and everybody drinks from equally, which is the diffuse field assumption, and Rindel names it as a limitation in his own conclusion.

But you can hear the next table. Not the room's memory of the next table, the actual person, arriving directly across two metres of air. So the obvious thing to do is put the people somewhere and add that term back, which is the textbook direct-plus-diffuse room equation and not an invention:

Lp = LW + 10 log10( Q / 4πr2 + 4/A )

Then close the Lombard loop for each talker separately, so each person answers the noise they are personally sitting in rather than a room average. With the direct term switched off this reproduces Rindel's closed form to ten decimal places, which is the control. With it switched on, in the five rooms below, the room comes out between 1.8 and 3.3 decibels louder than the closed form says.

Instrument five · what the diffuse field leaves out

Diffuse field promises
77.2 dB
With neighbours audible
79.1 dB
Return on next doubling
-5.5 dB
Critical distance
1.7 m

The straight line is the promise: six decibels for every doubling of absorption, for ever. The curve that bends away from it is what happens when the people have positions. Absorption drains the reverberant field, and it does absolutely nothing to the sound arriving directly from a person two metres away, so as the treatment improves, the fraction of the noise that treatment can reach gets smaller. The rule decays into its own success.

How much does this matter in a room you could actually build? The sweep above is the bistro, twelve metres by eight by four. It has 352 square metres of surface. Even if every wall, the floor and the ceiling were perfect absorbers, which nothing is, the absorption area could not exceed about 375 square metres, and the last doubling available to it would return 3.38 decibels rather than 6.02. The point where the return has fallen to half sits at 453 square metres, which is to say: past the end of the room. You cannot treat your way out of your neighbours.

The size of the effect is one number, and it is pure geometry. Call ρ the direct sound energy the neighbours deliver as a fraction of the reverberant energy the same crowd puts into the room, computed from the seating plan alone. Then the excess is

excess = 1/(1 − c) × 10 log10(1 + ρ)  dB

which lands within 0.03 decibels of the simulated value in all five rooms. That factor 1/(1 − c) is the same loop gain again, doing the same job for a third time: the crowd amplifies its own geometry exactly as hard as it amplifies everything else. Checked directly at four slopes, the excess scales with the loop gain to within 3 percent.

RoomA, m2TalkersρDiffuse PlacedExcessPredicted

Over forty independently generated seating plans in the same room, the excess is 1.67 decibels with a standard deviation of 0.16, and it is positive in every single plan. It is a property of the room, not of where the chairs happened to fall.

The apparatus

What is measured, and by whom

What is modelled, and therefore uncertain

Reproduce it

Two scripts, no dependencies. The first re-derives every published number in Rindel 2010 from the equations and compares them with the printed tables. The second runs the spatial experiment and writes the artifact this page reads.

node research/room-lombard/reproduce-rindel.mjs node research/room-lombard/spatial-sim.mjs

The page itself is checked end to end by verify-nobody-chose-this-volume.mjs, which drives it in a real browser and confirms that the numbers rendered here are the ones the lab notebook computes.

Sources

Étienne Lombard's original report is Le signe de l'élévation de la voix, Annales des Maladies de l'Oreille et du Larynx 37 (1911) 101–119. It is cited here from the secondary literature above; we have not read the original. Worth noting while we are being careful: Rindel's prose in both papers says Lombard first reported the effect “as early as 1909”, while the reference each of them gives is the 1911 paper. We follow the reference rather than the sentence, and flag the gap rather than quietly pick a side.