You Were Already Answering
Somebody asks you a question. You answer about a fifth of a second later. But it takes your brain roughly six hundred milliseconds to turn a thought into a single spoken word, and a second and a half to build a sentence. The arithmetic does not close. The only way out is that you were assembling the reply while they were still talking, and you launched it on a guess about when they would stop.
This page asks you to prove that on yourself. Not by believing the numbers below, but by pressing a key at the end of a sentence and discovering that your finger moved before you could possibly have known.
The two numbers that do not fit
Both numbers are well measured and they are incompatible with the obvious story. If you waited for the end of my turn, then started building your reply, your first syllable would arrive around 600 milliseconds later at the very best. It does not. It arrives in about 200, and very often it arrives early, on top of my last word.
So comprehension and production must overlap. You are listening and composing at the same time, and then you are timing your entry to a moment that has not happened yet. Conversation is not a game of catch. It is two people finishing each other's turns with a precision neither of them notices.
Test one: your own reaction floor, then your own projection
Two short blocks. First we measure how fast you can respond to something you could not anticipate. That is your floor, and nothing you do afterwards can beat it by reacting. Then sentences appear one word at a time, and your job is to press at the instant the sentence ends. Some of them arrive as words. Some arrive as blocks with the same rhythm and the same lengths, but nothing to read.
The projection test
Ready
Eleven trials. Three to find your floor, eight to beat it.
Press SPACE, or tap anywhere in the dark panel. Sound is not used anywhere on this page. The words arrive at roughly 210 a minute, a brisk conversational pace, and each sentence closes on a short word that is gone again in under 300 milliseconds. That last detail is the whole point: it is shorter than most people's reaction time, so landing on it means you did not wait to see it.
What you just did
| Condition | trials | median | launched too early to be a reaction |
|---|
Your run is four trials per condition, which is a demonstration, not an experiment. It can show you the effect on yourself; it cannot establish the effect in general. What establishes it in general is de Ruiter, Mitterer and Enfield (2006), who ran the same logic on real recorded speech with real participants, and found that stripping the intonation out of a turn did not hurt people's ability to time its end, while stripping out the words did. It is the language, not the tune, that tells you when someone is about to stop.
Test two: the gap, measured here
The 200 millisecond figure is quoted so often that it is worth measuring rather than repeating. Below is the gap between turns computed from scratch out of hours of recorded conversation: meetings, participants, every one of them wearing their own microphone so that their speech is timed independently of everyone else's. Negative means the next speaker started before the last one stopped.
Floor transfer offsets, AMI Meeting Corpus
Drag it to 600 milliseconds, the time it takes to get one word out. Most of the conversation is already to the left of that line. Those replies cannot have been started when the previous speaker stopped, because there was not enough time left to build them.
Two honest qualifications, because this corpus is not the one the textbook number comes from. These are four-person meetings, and meetings are combative in a way that two people on a telephone are not: people cut in, and the median here lands at , on the overlapping side of zero, where the dyadic literature reports something closer to +170. And the peak of the distribution sits at , which is the more interesting fact anyway. The most common thing that happens between two turns in a real conversation is not a pause. It is a dead heat.
Test three: the silence that answers for you
The gap is not only fast, it is meaningful. A delay is itself a message, and everybody in the room can read it. Drag the pause and see what has actually been measured about silences of that length.
How long before the silence starts talking
This is why the pause before a favour is refused feels so loud, and why the pause before it is granted does not exist. You have been reading that signal your whole life at a resolution of a couple of hundred milliseconds, and you have never once had to think about it.
The check
Every number on this page is either recomputed from a corpus in this repository or attributed to a named published source, and the two are kept apart.
Measured here. The distribution above is computed by research/turn-taking-gap/measure.mjs from the AMI Meeting Corpus manual annotations v1.6.2 (CC BY 4.0). floor transfers over hours. Median , mode , mean . of transfers complete inside 600 ms. The mean is printed beside the median deliberately: it is the statistic that screams first when a definition is wrong, and it screamed three times during this build.
The one free parameter, shown moving. Utterances are built by merging a speaker's segments across silences shorter than P. Sweeping P from 50 ms to 1000 ms moves the median by and the overlap rate by under a percentage point, so the result does not rest on that choice.
Taken from published work, not measured here. The ~200 ms cross-linguistic gap and the zero modal offset are Stivers et al. (2009). The ~600 ms single word and ~1,500 ms sentence latencies are Indefrey & Levelt (2004) and Griffin & Bock (2000), as compiled by Levinson & Torreira (2015). The words-not-tune result is de Ruiter, Mitterer & Enfield (2006). The 700 ms preference threshold is Kendrick & Torreira (2015); the brain response to delayed answers is Bögels, Kendrick & Levinson (2015). Full list below.
The test's two free choices, named. A reaction faster than 100 ms is discarded rather than counted, because nobody responds to a visual signal that quickly and keeping it would hand you an impossibly low floor and then flatter you against it. And a press only counts as evidence if it landed within 400 ms of the true ending, so that hammering the key early scores nothing. Both numbers are stated here rather than buried because both of them move the result.
Your own numbers are computed in your browser, from your own key presses, and are never sent anywhere. Nothing on this page makes a network request.
What this page cannot tell you, and one corpus that could not either
The corpus is meetings, not chat. AMI is four people around a table, most of them working through a design task, many of them not native English speakers. The overlap rate here ( of transfers begin before the previous speaker stops) is far higher than the roughly 30% reported for two-party telephone conversation, and that difference is probably real rather than a mistake: four people competing for one floor behave differently from two. The median here is negative. The published dyadic median is positive. Both can be true.
The test above is a visual analogue, not a replication. Real turn-end projection happens in sound, over speech, in real time. Reading words as they appear is a different task, and the numbers you get here are not comparable to published reaction times. What it can do honestly is separate two things on you personally: whether your press could have been a reaction to the last word appearing, or must have been launched before it.
A corpus we tried first and had to throw out. The obvious free corpus of American conversation is the Santa Barbara Corpus, and it cannot answer this question at all. Its transcript timestamps are largely chained: one unit's end time is the next unit's start time, 89% of the time within a speaker and 41% of the time across a speaker change. So 41% of its speaker transitions have a gap of exactly 0.000 seconds, not because people are that fast but because nobody measured it separately. Run research/turn-taking-gap/sbc-check.mjs and it prints the whole autopsy. AMI's own forced-aligned word timings fail the same test even harder (91.5% chained within a speaker), which is why this page uses its hand-marked segment boundaries instead. That choice is not cosmetic: running the identical pipeline over the word times instead moves the median by and the peak by , which is larger than the quantity being measured. Same code, same corpus, same hour of conversation, different column of timestamps. The only defence against picking the wrong one is to run the test, and the reason this paragraph exists is that the first two versions of this page did not.
Sources
- Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E. & Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. PNAS 106(26), 10587-10592. Ten languages, 101 conversations. Modal response offset 0 ms; cross-language median +100 ms; language means from +7 ms (Japanese) to +469 ms (Danish), all within about 250 ms of the cross-language mean. doi:10.1073/pnas.0903616106
- Levinson, S. C. & Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology 6:731. The review that states the puzzle in one line: gaps are of the order of 200 ms, production latencies are over 600 ms, so turn ends must be predicted. doi:10.3389/fpsyg.2015.00731
- Roberts, S. G., Torreira, F. & Levinson, S. C. (2015). The effects of processing and sequence organization on the timing of turn taking. Frontiers in Psychology 6:509. 19,754 transitions from 348 conversations: mean 187 ms, median 168 ms, mode 169 ms. Negative answers arrive 55 ms later on average than positive ones. doi:10.3389/fpsyg.2015.00509
- de Ruiter, J. P., Mitterer, H. & Enfield, N. J. (2006). Projecting the end of a speaker's turn: a cognitive cornerstone of conversation. Language 82(3), 515-535. Flattening the intonation left projection accuracy unchanged; removing the words by low-pass filtering wrecked it. Intonation is neither necessary nor sufficient. doi:10.1353/lan.2006.0130
- Bögels, S., Magyari, L. & Levinson, S. C. (2015). Neural signatures of response planning occur midway through an incoming question in conversation. Scientific Reports 5:12881. EEG shows response planning beginning within half a second of the point where the answer becomes retrievable, in some items seconds before the question ends. doi:10.1038/srep12881
- Kendrick, K. H. & Torreira, F. (2015). The timing and construction of preference: a quantitative study. Discourse Processes 52(4), 255-289. Past roughly 700 ms of delay, dispreferred responses outnumber preferred ones. doi:10.1080/0163853X.2014.955997
- Bögels, S., Kendrick, K. H. & Levinson, S. C. (2015). Never say no: how the brain interprets the pregnant pause in conversation. PLOS ONE 10(12):e0145474. After a 300 ms gap a "no" produces an N400, the brain's signature of an unexpected meaning. After 1000 ms it does not: the silence has already prepared the listener. doi:10.1371/journal.pone.0145474
- Indefrey, P. & Levelt, W. J. M. (2004). The spatial and temporal signatures of word production components. Cognition 92(1-2), 101-144. The ~600 ms estimate for producing a single word from a concept.
- Griffin, Z. M. & Bock, K. (2000). What the eyes say about speaking. Psychological Science 11(4), 274-279. Sentence production latencies of roughly 1,500 ms.
- Carletta, J. et al. (2006). The AMI Meeting Corpus: a pre-announcement. MLMI 2005, LNCS 3869, 28-39. Manual annotations v1.6.2, CC BY 4.0. The corpus measured on this page. groups.inf.ed.ac.uk/ami/corpus
- Du Bois, J. W., Chafe, W. L., Meyer, C., Thompson, S. A., Englebretson, R. & Martey, N. (2000-2005). Santa Barbara Corpus of Spoken American English, parts 1-4. CC BY-ND 3.0 US. Downloaded, tested, and reported here as unusable for this particular measurement, with the test shown. linguistics.ucsb.edu