Nobody Chose This Volume
You are in a full restaurant, leaning across the table, asking someone to repeat themselves for the third time. Nobody in the room is shouting on purpose. Every single person is raising their voice by exactly the amount that seems reasonable, and the room has settled at a level that no one in it would have picked.
Two facts, both measured, sit awkwardly together.
The first: put a person in noise and they get louder. Étienne Lombard reported it in 1911, and it is not a decision, it is a reflex. You do it while insisting you are not doing it. The rate is about half a decibel of extra voice for every decibel of noise.
The second: sound adds in energy, not in loudness. Two people talking make about three decibels more noise than one, not twice as much. Double the crowd, add three decibels. That is the arithmetic of incoherent sources and it does not care how you feel about it.
Hold those two together and the second one turns out to be wrong, in a room with people in it. Doubling the crowd does not add three decibels. It adds six. The extra three come from everybody answering the first three, and then answering each other's answers, all the way down. The room is an amplifier, and the gain is exactly two.
The loop, drawn
Here is the whole mechanism as two straight lines.
The rising line is what people do: the level you speak at, given the noise you are sitting in. It starts at 55 decibels when the room is at 45, which is where the reflex switches on, and it climbs with slope c, the Lombard slope. The other line is what the room does: the noise the room holds, given the level everyone is speaking at. Its slope is exactly 1, because a decibel in is a decibel out.
The room settles where the lines cross. Drag the slope.
Instrument one · the fixed point
Two things are worth staring at.
At c = 0 the lines cross low and the loop gain is 1: people would simply talk, and the room would be whatever the arithmetic of adding voices says it is. Every decibel of gain above that is the crowd answering itself.
Push c toward 1 and the lines become parallel. There is no crossing point. That is the runaway that people mean when they say a party gets out of hand, and the reason it does not actually happen is the whole of the reassurance available here: the measured slope is about a half, so the lines do cross, so the room does settle. It settles somewhere unpleasant, but it settles.
A room, and what it does to you
Now put numbers on it. Below is J. H. Rindel's prediction model from 2010, which is the above two lines solved for their crossing point and nothing more:
where A is the room's absorption area in square metres and NS is the number of people actually talking at any given moment, which is fewer than the number of people present, because at a table of four only one person talks at a time. Rindel calls the ratio the group size. At c = 0.5 the whole thing collapses to something you can almost do in your head:
Build a room. The absorption comes from the Sabine equation, so the two things you set are the volume and how long a clap takes to die away in the empty room.
Instrument two · build the room, then listen to it
Synthesised, not recorded: speech-shaped noise with syllable-rate envelopes, one voice per talker up to twelve and a steady bed for the rest. Your volume knob sets the absolute level, so treat the loudness as arbitrary. What is not arbitrary is the ratio: the near voice and the room are mixed at exactly the signal-to-noise ratio the model computes above. Turning the reflex off recomputes the room at c = 0, which is the same crowd not answering itself.
Two of those readouts deserve a word.
Signal to noise is the level of the person you came with, measured a metre away, minus the level of the room. Below about −3 decibels, the acoustics literature stops calling conversation sufficient. Note where that lands relative to the rooms you actually eat in.
Acoustic capacity is Rindel's proposal that a room should be labelled with the largest number of people at which talking still works, the way it is labelled with a fire limit. The rule of thumb is memorable:
Volume in cubic metres, reverberation time in seconds. It is a simplification, and it is worth knowing what of. Solve the signal-to-noise condition exactly, with the diners' own bodies counted as absorbers, and the divisor comes out at 20.29, not a round 20, for the group size and per-person absorption that paper assumes. Change the assumption about how talkative the room is and the divisor moves a long way: . The formula is honest; it is the group size that is soft.
The doubling that pays twice
Here is the part that is genuinely useful, and it falls out of the same factor of two.
Take an ordinary noise source, a ventilation fan say, and double a room's absorption. The fan pushes out the same power as before, the room holds less of it, and the level drops by three decibels. That is the textbook result and everyone in building acoustics knows it.
Now do it in a room full of people. The level drops by three decibels, so everyone stops straining, so everyone drops their voice by half of that, so the level drops again, and the whole series sums to six.
Instrument three · the same treatment, two rooms
The same number, 1/(1 − c), governs the disease and the cure. It is why a crowded room is worse than you would predict, and it is why hanging absorption in one is better than you would predict. Acoustic treatment works on people better than it works on machines, because machines do not have opinions about whether they are being heard.
Ten real rooms
Hodgson, Steiniger and Razavi measured ten eating places in Vancouver in 2007, and Rindel tabulated them against his model. The table below is not copied from the paper. It is computed in your browser from the volume, the empty reverberation time and the number of seats, using the two equations above and nothing else, and then compared with what was actually measured with a sound level meter in the room.
Instrument four · the model against the meter
| Room | V, m3 | T, s | Seats | A, m2 | Predicted | Measured | Miss | Capacity | Conversation |
|---|
C: cafeteria. B: bistro. R: restaurant. S: senior residence dining room. Group size 4 for the first eight, 8 for the last two, and 0.5 m2 of absorption per person, exactly as the paper assumes.
Two things fall out of that table that are worth saying plainly.
The first is that the model works. Eight of the ten land within about a decibel, from a calculation with three inputs, against a real meter in a real room. Predicting the noise of a crowd from the shape of the room is not supposed to be that easy.
The second is the capacity column. Filled to their own seat count, eight of these ten rooms are past the point where the literature calls conversation sufficient. The two that are not are the dining rooms in the senior residences, and they pass for a reason that is not architectural: the paper needed a group size of eight to fit them, against three or four everywhere else. Fewer people talking at once. It is the quietest room in the study and the loneliest number in the table.
Where the model stops being true
The equation above contains no distances. Not one. It treats the room as a single tank of sound that every talker pours into and everybody drinks from equally, which is the diffuse field assumption, and Rindel names it as a limitation in his own conclusion.
But you can hear the next table. Not the room's memory of the next table, the actual person, arriving directly across two metres of air. So the obvious thing to do is put the people somewhere and add that term back, which is the textbook direct-plus-diffuse room equation and not an invention:
Then close the Lombard loop for each talker separately, so each person answers the noise they are personally sitting in rather than a room average. With the direct term switched off this reproduces Rindel's closed form to ten decimal places, which is the control. With it switched on, in the five rooms below, the room comes out between 1.8 and 3.3 decibels louder than the closed form says.
Instrument five · what the diffuse field leaves out
The straight line is the promise: six decibels for every doubling of absorption, for ever. The curve that bends away from it is what happens when the people have positions. Absorption drains the reverberant field, and it does absolutely nothing to the sound arriving directly from a person two metres away, so as the treatment improves, the fraction of the noise that treatment can reach gets smaller. The rule decays into its own success.
How much does this matter in a room you could actually build? The sweep above is the bistro, twelve metres by eight by four. It has 352 square metres of surface. Even if every wall, the floor and the ceiling were perfect absorbers, which nothing is, the absorption area could not exceed about 375 square metres, and the last doubling available to it would return 3.38 decibels rather than 6.02. The point where the return has fallen to half sits at 453 square metres, which is to say: past the end of the room. You cannot treat your way out of your neighbours.
The size of the effect is one number, and it is pure geometry. Call ρ the direct sound energy the neighbours deliver as a fraction of the reverberant energy the same crowd puts into the room, computed from the seating plan alone. Then the excess is
which lands within 0.03 decibels of the simulated value in all five rooms. That factor 1/(1 − c) is the same loop gain again, doing the same job for a third time: the crowd amplifies its own geometry exactly as hard as it amplifies everything else. Checked directly at four slopes, the excess scales with the loop gain to within 3 percent.
| Room | A, m2 | Talkers | ρ | Diffuse | Placed | Excess | Predicted |
|---|
Over forty independently generated seating plans in the same room, the excess is 1.67 decibels with a standard deviation of 0.16, and it is positive in every single plan. It is a property of the room, not of where the chairs happened to fall.
The apparatus
What is measured, and by whom
- The Lombard slope, c = 0.5 dB/dB. Lazarus's 1986 review put the range at 0.5 to 0.7. Hodgson and colleagues fitted 0.69 to their ten rooms. Rindel tried the range against three independent datasets and found that only 0.5 fits: at 0.6 the model misses by 5 to 12 decibels and at 0.7 by 14 to 27. That is a live disagreement in the literature, not a settled constant, and everything on this page inherits it. At 0.69 the loop gain would be 3.2 and a doubling of the crowd would cost 9.7 decibels rather than 6.0.
- The reflex threshold, 45 dB. Below that ambient the effect does not operate and the model does not apply.
- The validity band. Rindel states the model holds for speech between 55 and 75 decibels. Converted through the signal-to-noise equation, that is an absorption-per-talker range of 2.51 to 251 square metres, exactly two decades, and the instruments above stay inside it.
- The ten rooms were measured by Hodgson, Steiniger and Razavi and reported with a range of levels through a day of normal operation; the comparison uses the highest level in each reported range, as Rindel did.
What is modelled, and therefore uncertain
- Group size is the soft parameter and it dominates. It is a social fact wearing an acoustic costume: how many people are in the room per person actually talking. The paper's own fits range from 2.5 in the liveliest bistro to 9 in a senior residence. Moving it from 2.5 to 5 moves the predicted level by about six decibels. Any number this page gives you inherits that.
- Talker directivity is taken as Q = 2, a hemisphere, which is Rindel's approximation. Real speech is more forward-directed than that and frequency-dependent.
- The spatial simulation adds the neighbours' direct sound and nothing else. It does not model shielding by furniture or bodies, non-diffuse decay in flat low rooms, early reflections, or the fact that people turn their heads. It is one correction, isolated on purpose, not a room acoustics package.
- The seating plans are randomly generated by rejection sampling at a stated minimum table pitch, not surveys of real floor plans.
- The audio is synthesised from filtered noise. The ratio between the near voice and the room is the model's; the absolute level is your volume control and means nothing.
- None of this is a measurement of any actual restaurant by us. We did not take a sound level meter anywhere. Every measured number on this page belongs to the papers cited below.
Reproduce it
Two scripts, no dependencies. The first re-derives every published number in Rindel 2010 from the equations and compares them with the printed tables. The second runs the spatial experiment and writes the artifact this page reads.
The page itself is checked end to end by verify-nobody-chose-this-volume.mjs, which drives it in a real browser and confirms that the numbers rendered here are the ones the lab notebook computes.
Sources
- J. H. Rindel, Verbal communication and noise in eating establishments, Applied Acoustics 71 (2010) 1156–1161. doi:10.1016/j.apacoust.2010.07.005. The prediction model, the ten-room comparison, and the quality-of-communication table.
- J. H. Rindel, The acoustics of places for social gatherings, plenary lecture, EuroNoise 2015, Maastricht, 2429–2434. The acoustic capacity concept and the V/(20T) rule.
- M. Hodgson, G. Steiniger, Z. Razavi, Measurement and prediction of speech and noise levels and the Lombard effect in eating establishments, Journal of the Acoustical Society of America 121 (2007) 2023–2033. The ten rooms, and the fitted slope of 0.69 dB/dB.
- H. Lazarus, Prediction of verbal communication in noise, a review: part 1, Applied Acoustics 19 (1986) 439–464, and part 2, Applied Acoustics 20 (1987) 245–261. The Lombard slope range and the quality-of-communication labels.
- W. R. MacLean, On the acoustics of cocktail parties, Journal of the Acoustical Society of America 31 (1959) 79–80. The critical-number argument, cited here as the origin of the runaway idea and not as a description of what the measurements show.
- S. K. Tang, D. W. T. Chan, K. C. Chan, Prediction of sound-pressure level in an occupied enclosure, JASA 101 (1997) 2990–2993, and M. P. N. Navarro, R. L. Pimentel, Speech interference in food courts of shopping centres, Applied Acoustics 68 (2007) 364–375. The canteen and food court datasets Rindel fits.
- ISO 9921:2003, Ergonomics, assessment of speech communication. The vocal effort scale in six-decibel steps.
Étienne Lombard's original report is Le signe de l'élévation de la voix, Annales des Maladies de l'Oreille et du Larynx 37 (1911) 101–119. It is cited here from the secondary literature above; we have not read the original. Worth noting while we are being careful: Rindel's prose in both papers says Lombard first reported the effect “as early as 1909”, while the reference each of them gives is the 1911 paper. We follow the reference rather than the sentence, and flag the gap rather than quietly pick a side.