Artificial Wasteland · the mind seam

Nothing Funny Happened

Almost nobody laughs at jokes. Here is what people were really saying, verbatim, in the moment before they laughed, taken from two open corpora of recorded conversation. See whether you can tell a line that got a laugh from one that did not. Then watch a famous finding come apart: laughter is supposed to almost never interrupt a phrase, seven times in a thousand, and in both corpora it does it about ten times in a hundred.

Somebody laughs roughly every couple of minutes that they are talking to another person. Ask what they are laughing at and the obvious answer is: something funny. That answer is wrong, and it has been known to be wrong since Robert Provine and three undergraduates stood around in shopping malls in the early nineties writing down what people said just before they laughed. Of laughs they logged, the ones that followed anything like an attempt at humour were a minority, and the rest followed remarks like it was nice meeting you too.

This page is not a retelling of that. Two corpora of real recorded conversation are now published under open licences with every laugh marked by hand, which means the counting can be done again, by anyone, on material nobody chose for being funny. That is laughs in hours of recorded meetings, and more in hours of American dinner tables, phone calls, a card game and a sermon. Everything below is measured from those, and every script is in the open. Some of it confirms Provine. One large piece of it does not.

I. What they were actually saying

Below are two things one person said, in the same recorded meeting, the same number of words long. One of them ran straight into that person laughing. The other did not, and nobody in the room laughed within two seconds of it either. Pick the one that got the laugh.

Which one was followed by laughter?

0 of 0

Drawn with a fixed seed from matched pairs, not chosen by hand. Verbatim from the AMI Meeting Corpus, CC BY 4.0. Chance is 50%.

There is no aggregate score to compare yours against, because we have not run this on a crowd. What we can do is show what the pool looks like. Here is a random sample of the words that ran directly into a laugh, in the same breath, by the person who then laughed.

Said, and then laughed

Verbatim, AMI Meeting Corpus (CC BY 4.0), sampled with a fixed seed from laughs that had between four and sixteen of the laugher's own words in front of them.

The single commonest word in the English language immediately before somebody laughs is not a punchline. Across laughs with one of the laugher's own words in front of it, here is what that word was, and how much commoner it is in that position than in the corpus at large.

Last word before the laughcountshare hereshare overalllift

Yeah. Okay. Uh. Yes. Right. These are the sounds of a conversation continuing, not of a joke landing. Provine put the share of pre-laugh comments that were "remotely humorous" at to percent in his later papers, and at under percent in an earlier one. We have not scored humour here, because we have no way to do that honestly without a panel of judges. What we can do is show you the material and let you decide, which is what the game above is for.

II. Where the laugh lands

The second and stranger claim is about placement. Provine's finding was that laughter does not cut into the structure of a sentence: it arrives where a comma or a full stop would, at the end of a phrase. A speaker will say you are going where? ha-ha, and essentially never you are going, ha-ha, where? In his sample the exception rate was laughs out of . He called it the punctuation effect.

Below is a real utterance from a real meeting, with the laugh taken out. Click the gap where you think it went.

Put the laugh back

0 of 0

Verbatim AMI utterances that contain a laugh with the speaker's own words on both sides of it. Punctuation is the transcriber's.

Now the counting. Of the laughs in the meeting corpus, are an utterance all by themselves, with no words attached. Of the that do sit in an utterance with words, most are at one end of it. That leaves laughs with the laugher's own words on both sides, which are the only ones that can test the claim at all, because everywhere else the answer is forced.

Of those : arrive straight after a punctuation mark, and split a phrase.

The null is not a guess. Each laugh was reassigned at random, ten thousand times, among the interior slots of its own utterance, which holds utterance length fixed and destroys only the choice of where inside it to laugh. That gives an expected boundary rate of percent. The observed rate is percent. The best of the random runs reached percent, and of them reached the observed value.

So there is an effect, and it is large. But it is nothing like seven in a thousand. Two out of five interior laughs cut into a phrase.

Is that just the transcriber's comma?

It could be. The punctuation in that corpus was typed by a person who could hear the laugh, and a laugh is exactly the sort of thing that might make somebody reach for a full stop. If so, the correlation would be between one transcriber's two decisions and not between laughter and grammar at all.

So the punctuation was thrown away. The bare word sequence of each of those utterances went to an off-the-shelf English dependency parser () that never saw the laugh and never saw a comma, and the question became purely structural: how many syntactic arcs cross the point where the laugh happened? At a clean break between phrases, few do. In the middle of a noun phrase, several do. Every utterance is its own control, compared against every other position inside it.

Dependency arcs crossing the laugh's position

Scored on utterances. Mean percentile rank of the laugh's position among the alternatives in its own utterance: , where 0.5 would be chance and 1.0 would be the cleanest break every time. Over permutations, came out at or below the observed mean.

The effect is real and it is not an artefact of the comma. A parser that has no idea a laugh happened puts the laugh's position at a cleaner syntactic seam than the rest of the same sentence, by a wide margin. Better still: the laughs the transcriber did not put a comma in front of, the ones scored as splitting a phrase, still sit at crossing arcs against for the rest of their own utterances. Even the violations are at soft seams. The effect is graded, not a rule.

III. What the instrument cannot see

Here is where the honest version of this page diverges from the tidy one.

The obvious way to break the punctuation effect is to laugh while still saying the words. Linguists call it speech-laughter, and you do it constantly: she is his long-term heh friend. Tian, Mazzocconi and Ginzburg noticed in 2016 that Provine had excluded speech-laughs from his samples, in their words, "without any justification". Truong and Trouvain reported that in the very meeting corpora everyone uses, speech-laughs are sometimes ignored and sometimes inconsistently labelled, and judged that re-annotation was necessary. Gilmartin and colleagues put it flatly: these corpora have no method for annotating laughter that co-occurs with speech.

If that is true of the corpus this page has been counting, then the punctuation effect measured above is partly a property of the annotation scheme. A corpus that cannot write down a laugh inside a word will report that laughter never lands inside a word. So rather than take the warning on trust, we tested it.

Does a laugh ever overlap the laugher's own words?

Of laughs in the meeting corpus that carry a usable start and end time, the number whose interval overlaps one of that same speaker's own word intervals is .

Two. Not two percent. Two laughs.

And it gets sharper. A laugh in that corpus can be entered as a marker in the word order with no duration at all: start time equal to end time. Where do those fall?

Where the laugh sitslaughsno duration recordedshare

Every single one of the interior laughs, the ones this whole section is built on, was entered without a duration. When a laugh lands in clear air the transcriber times it. When it lands inside somebody's sentence, the scheme records that it happened and where in the word order, but not when, and it never records it overlapping the words at all.

That does not invalidate the placement result: a human being deliberately put that marker between those two words, and that is a real judgement about where the laugh went. But it does mean the count of phrase-splitting laughs is a floor. The laughs most likely to break the rule are precisely the ones this instrument is worst at seeing, and Provine's own method, which excluded them on purpose, had the same blind spot in a stronger form. Two instruments, thirty years apart, blind in the same place.

IV. The number, from two corpora that share nothing

Which is why the second corpus matters. The Santa Barbara Corpus of Spoken American English is not a meeting corpus. It is recordings of Americans talking to each other, transcribed by a different team under a different scheme, in units marked from the rise and fall of the voice rather than from grammar. And its notation can do the thing the meeting corpus cannot: it writes laughter as the symbol @, one for each pulse, placed exactly where the pulse happened, including hard against a word in the middle of a phrase.

In that corpus, percent of laughs occupy a unit of speech entirely by themselves. For comparison, only percent of the units that contain words are a single word standing alone. Laughter really does get its own slot in the stream in a way that words do not.

But the statistic that answers Provine directly is the share of all laughs that land between two words. Here it is, in both corpora, against the published figure.

Nine point four percent and nine point six percent. Two corpora with no recordings, no speakers, no transcribers, no register and no notation in common, agreeing to within a fifth of a percentage point, and both about fourteen times the published rate. They are also in the same neighbourhood as the three other groups who have counted this on conversational data: Tian and colleagues found percent in French dyads and percent in Chinese; Maraev and colleagues, pooling laughs across five corpora, found percent as stand-alone bouts and percent as speech-laughs.

The punctuation effect is real. Laughter is not sprinkled over speech at random, and a parser that knows nothing about it will tell you so. But "over 99 percent" is not a fact about laughter. It is a fact about a method that could not see the exceptions, and it has been repeated for thirty years.

Provine himself has quietly reported a much weaker version. In a 2006 study of laughter among deaf signers, where laughter cannot be competing with speech for the vocal tract at all, the figure given is that laughter occurred about times more often at pauses and phrase boundaries than simultaneously with a signed utterance. That ratio is roughly three to one. Not a thousand to seven.

V. The person laughing is the person talking

Provine's other counterintuitive result was that the speaker laughs more than the audience, by percent. The meeting corpus can test that directly, because it knows who was talking at every instant.

For each laugh, the floor-holder is whoever produced the most recent word before it. In a meeting of four, the floor-holder is one person out of four. If laughter were the audience's response to a speaker's wit, the speaker's share of laughs would sit below a quarter.

The floor-holder produces percent of all laughter while being, on average, one person in four. Per person, the one who is talking laughs about times as often as each of the people listening. Provine said 46 percent more. This is a great deal more than that.

A talkative person holds the floor more often and might simply laugh more, which would manufacture this result out of nothing. So the null was built to kill exactly that: every speaker's laugh times were rotated together by one random offset around the meeting clock, which preserves how much each person laughs, how much each person holds the floor, and the spacing between one person's laughs, and destroys only the alignment between them. Over rotations the mean speaker share was percent, the highest was percent, and reached the observed value.

VI. What a laugh sounds like

Six real laughs, pulled out of the meeting recordings by their annotated times, with nothing done to them but a change of loudness so that six microphone gains can be compared by ear. The grey shape under each is its own amplitude envelope.

In 1991 Provine and Yong published the first quantitative description of laughter's sound: notes of about milliseconds, separated by about to , falling in amplitude across the bout. Ten years later Bachorowski, Smoski and Owren analysed laughs from people and concluded almost the opposite: that laughter is variable and complex rather than stereotyped, with intervals around milliseconds, half of Provine's figure.

We measured laughs from the meeting corpus, fetched a second at a time straight out of the recordings, and ran exactly the same measurement on matched stretches of ordinary speech by the same people. The control matters more than anything else here, because ordinary speech already has a rhythm near five per second: the syllable rate. A laugh that came out at five per second would prove nothing on its own.

Envelope smoothinglaughter periodspeech periodlaughter regularityspeech regularity

Three results, and only one of them is the expected one.

And the pulse count. The Santa Barbara transcribers hand-counted the pulses of every laugh in that corpus, one @ each, giving pulses in laughs, a mean of . The acoustic peak-picking on the other corpus, from audio, with no human in the loop, gives a median of . Two entirely different methods, two corpora, and the answer is somewhere between two and four pulses. The mental image of a long rolling ha-ha-ha-ha-ha is mostly wrong. Most laughs are two or three pulses and are over in of a second.

The laugh machine

Now build one. This is a synthesiser, not a recording: a buzz at a chosen pitch pushed through two formant filters and chopped into pulses. It proves nothing on its own. What it is good for is hearing what the published numbers actually sound like, and finding the edges of the thing by walking off them.

Provine reported that the vowel does not vary inside a bout: ha-ha-ha or ho-ho-ho, never ha-ho-ha-ho. That claim is one both camps agree on, and the checkbox lets you break it.

The check

What this does not show