the ground / stratum · Life

The Drift That Has Nowhere to Go

CEILING 14.7 d AVERAGE GAP 7.4 d COUPLING IN THE MODEL zero BENCHES five

Five working benches. On the third one there is a dial marked true coupling. Drag it to the largest effect anyone has ever published and watch the classic statistic barely notice.

Two women live together for a year, and one month their periods start on the same day. It is a small, strange intimacy, and almost everyone who has lived it treats it as evidence of something. There is a name for the something: menstrual synchrony, proposed in 1971, taught for decades, repeated everywhere.

This page is not going to argue with the feeling. The feeling is real, it is extremely common, and it has a cause. The cause is just not the cycles finding each other, and the reason we can be fairly confident of that is a piece of arithmetic you can operate below rather than an appeal to anybody's authority.

The claim, exactly

In 1971 Martha McClintock published two pages in Nature called Menstrual Synchrony and Suppression. The subjects were 135 women aged 17 to 22, all living in one dormitory at a suburban women's college. Three times across the academic year each was asked when her last two periods had begun, which reconstructed onset dates from late September to early April.

For each pair of roommates, and for each pair who had both named the other as the person they saw most often, the paper compared how far apart their onsets fell in October with how far apart they fell in March. The pairs came closer together. Pairs assigned at random from the same dormitory did not: that comparison returned P ≤ 0.8. Women grouped only by where their rooms were did not either. For fifteen groups of close friends the paper plots the median distance from the group's own average onset, and across the year it fell from roughly six and a half days to roughly four and a half.

Two things about that paper are worth having straight, because the version everyone repeats has lost both of them.

McClintock did not claim to have found a pheromone. The paper says the mechanism “is a question which still remains open for speculation and investigation.” A pheromone enters as a parallel to a known effect in mice, and the paper does go on to say that “some additional pheromonal effect among individuals of the group of females would be necessary” and “perhaps at least one female pheromone affects the timing of other female menstrual cycles.” But it says perhaps, and it says the mechanism is open. What it concluded firmly was about the finding, not the cause: “the significant factor in synchrony, then, is that the individuals of the group spend time together,” and, in its last sentence, that “although this is a preliminary study, the evidence for synchrony and suppression of the menstrual cycle is quite strong.” She was not tentative about the effect. She was careful about the mechanism, and it is the mechanism that thirty years of retelling promoted from speculation to fact.

The paper contradicts itself twice in two pages. On the first page the roommate result is P ≤ 0.0007. On the second, in a passage arguing against a shared light cycle as the explanation, the same result is P ≤ 0.007. Table 2 gives the two male-exposure groups as N = 56 and N = 31; the running text gives them as N = 42 and N = 33. Neither was ever corrected, which tells you something about how closely the most famous claim in this field was read.

None of that settles anything by itself. A sloppy paper can still be right. So take the quantity every study of this reduces to, which is how many days apart two women's periods begin, and look at what that quantity is able to do.

The gap cannot keep growing. Push two cycles apart far enough and they come back around the other side.

If your period starts twenty days after hers, it also starts about nine days before her next one, and nine is the honest number. The largest the gap can ever be is half a cycle. Everything below follows from that ceiling, and the first bench is two clocks with no wire between them.

Bench one · two clocks, no wire

two years · onsets, bleeding, and the gap
29.5 d
28.0 d
3.9 d
nothing in this code lets either woman see the other

The ceiling, and the average it forces

Watch that gap long enough and you notice it does not wander freely. It is trapped between zero and half a cycle, and inside that box it has an average it keeps coming back to. For a cycle of about a month, the average is about a week.

The reasoning is short enough to check by hand. If two clocks have no relationship, the moment one of them strikes is a moment picked at random inside the other one's cycle. The distance from a random point to the nearest of a set of marks spaced L apart averages L/4. For a cycle near 29.3 days that is 7.33 days, and the second bench measures it across tens of thousands of pairs who have never met.

This is not a new observation and it is not ours. Strassmann set it out in the 1997 Dogon paper, crediting Wallis, her own earlier work and Wilson, and put it in one line again in 1999: for a 28-day cycle, taken there as an example rather than a rule, the most two women can be out of phase is fourteen days, on average the onsets will be seven days apart, and “fully half the time they should be even closer.” What this page adds is not the argument. It is the ability to put your hands on it.

Bench two · the shape of the box

every gap, across thousands of unrelated pairs
29.3 d

Why a study of this has to find something

Now put the ceiling together with the way the question gets asked. Every version of the classic design does the same thing: record how far apart two women's periods are near the start of the observation, record it again some cycles later, and report the change.

Consider what happens to a pair that starts unusually far apart. The gap has an average it returns to, so a pair sitting well above that average has only one place to go. It will close. Not because anything drew the women together, but because the number was high and the number has a mean.

That is regression to the mean, the plainest of all statistical traps, and on the third bench it is worth three and a half days of apparent convergence. The third bench runs the procedure on simulated women who provably cannot influence one another, because the code that generates them has no term that would let them.

Bench three · the procedure, run on strangers

one line per pair
0.0 d/cyc

That is the whole difficulty in one picture. The number the procedure reports hardly moves when the truth underneath it changes. With no coupling at all it reports 3.6 days of convergence among the pairs that started above the median, with a run to run scatter of 0.5 days. Turn the true coupling up to three days of pull per cycle, which is nearly twice that, and it reports 4.4.

A measurement that returns nearly the same answer whether or not the thing exists is not a measurement of that thing.

Be precise about the strength of that, because overstating it would be the same sin in the other direction. The classic statistic is not blind. It does creep upward, and by a pull of nine days per cycle it is plainly elevated. The claim is narrower and it is about the range that actually exists: set the dial to 1.7 days per cycle, which is the largest per cycle shift in cycle length we can find published from a human chemical signal, and the reported convergence moves by 0.45 of its own run to run scatter. You could not tell those two worlds apart with this instrument if you ran the study a hundred times. Over the same change, the coherence of the group moves 0.93 of its scatter, twice as far.

The second meter on that bench is the same simulation measured a different way, by how tightly the whole group’s phases cluster, and that one moves: from 0.18 with no coupling to 0.24 at three days per cycle, a separation of about two standard deviations where the gap statistic manages one and a half. Drag the dial and watch the two meters come apart. Everything about the world is the same for both of them. Only the question differs.

Somebody did check, in 1992

H. Clyde Wilson went through the synchrony literature and found that every study in it used McClintock's design. In Psychoneuroendocrinology in 1992 the review listed three errors built into that design, and the list is worth reading in his own words:

Three errors are inherent in research based on her model: (1) an implicit assumption that differences between menses onsets of randomly paired subjects vary randomly over consecutive onsets, (2) an incorrect procedure for determining the initial onset absolute difference between subjects, and (3) exclusion of subjects or some onsets of subjects who do not have the number of onsets specified by the research design. All of these errors increase the probability of finding menstrual synchrony in a sample.

The first is the one the benches above are about. The differences do not vary randomly over consecutive onsets: they live under a ceiling, with a mean they return to, so a pair measured when it is far apart is a pair about to look like it is converging. The second inflates the starting figure, which makes the same effect larger. The third quietly removes exactly the women whose cycles are too irregular to fit the design, which is to say the women who would have shown the artifact for what it is.

Wilson's own summary of what happens when they are corrected, across the studies reviewed: “no significant levels of menstrual synchrony occur when these errors are corrected. Menstrual synchrony is not demonstrated in any of the experiments or studies.” The review also noticed a pattern that reads, in hindsight, like a confession from the whole literature at once. The studies that failed to find synchrony reported trouble with subjects whose cycles were irregular. The studies that found it reported no such trouble.

How hard would a real pull have to be?

Suppose, generously, that there is a signal. Suppose one woman's period can shift another woman's next one, by some number of days, in whichever direction would bring them together. How strong would that shift have to be before two women actually locked on to each other?

This is an old question with a clean answer, and it is the same answer that governs pendulum clocks on a shared wall, heart pacemaker cells, and the fireflies that flash in unison. Two oscillators pulling on each other lock when the coupling beats the mismatch in their natural rates. Below that, they slip past one another forever. For the mutual case, where each is pulling on the other, the threshold works out at half the mismatch; a single oscillator driven by an external clock it cannot affect needs the full mismatch instead, and the pair is the case that matters here.

Two women whose average cycles differ by two days are mismatched by two days per cycle. With nothing else going on, a pull of 1.0 day per cycle from each of them is exactly enough, which is the textbook boundary for two mutually coupled oscillators, half the mismatch, and which the bench recovers on its own without being told it. But there is something else going on, and it is large.

Bench four · the pull you would need

smallest pull that holds a pair together
3.9 d

Here is the part that decides it, and it has nothing to do with whether humans have pheromones at all. A woman's own cycles vary. Not by a rounding error: by days. In the Apple Women's Health Study, across 165,668 cycles, the within person standard deviation of cycle length was 3.79 days even at the most regular ages, and above five days in the teens and the late forties.

Put two women with identical average cycle lengths on exactly the same start date, with no coupling whatsoever, and one month later they are 4.39 days apart on average, purely from that variation. The closed form is her standard deviation times the square root of two times the square root of two over pi, which comes to 4.40. Six months of it and they are as far apart as two strangers.

Her own clock rattles harder than any pull that has ever been measured on it.

That is the number to hold on to, and the comparison it has to be made against needs saying carefully, because this is the one place where a convenient superlative could slip past unchecked. The largest shift in cycle length we can find reported from a human chemical signal is 1.7 days, from Stern and McClintock in 1998, and it is a one off shift, not a sustained pull toward a partner. Against her own month to month variation, the fourth bench says that holding a typical pair together would take about 2.8 days per cycle, every cycle, aimed correctly each time.

What happened when people looked properly

The strongest test of this was not run in a dormitory. Beverly Strassmann worked with the Dogon of Mali, where menstruating women stay in a menstrual hut, so attendance could be counted directly instead of remembered and reported. The women at the two huts in the study village were counted on each of 736 consecutive days, and the record was cross checked against urinary hormone assays. That gave 477 complete cycles from 58 women, recorded by observation rather than recall.

If menstrual onsets clustered, the number of days carrying two or three onsets would exceed what independence predicts. Strassmann's Table 3, and our arithmetic on it:

The conclusion, in the paper's own words: “the null hypothesis that the women's menstrual onsets were independent cannot be rejected.” And, on what that ought to do to the debate: “it shifts the burden of proof onto those who argue that the phenomenon exists.”

The rest of the record points the same way. Twenty nine cohabiting lesbian couples, keeping daily prospective records over three cycles, showed no convergence and mostly diverged (Trevathan and colleagues, 1993). Eighteen pairs and twenty one triples of Polish students across five months: nothing, and social contact was unrelated to the onset difference (Ziomkiewicz, 2006). A hundred and eighty six Chinese women in dormitories tracked for over a year: nothing, and on re examination the group synchrony in the 1971 study was “at the level of chance” (Yang and Schank, 2006). Reviewing the field in 2013, Harris and Vitzthum wrote that apparent synchronisation “is readily attributable to chance convergence arising from the finite and variable length of menstrual cycles and the rules of probability,” and that the interesting question by now is why so few people doubt it.

One widely shared piece of evidence deserves a warning label rather than a citation. In 2017 the cycle tracking app Clue reported that of 360 pairs of users, 273 had drifted further apart and 79 had come closer. That is the right direction for everything above, and we are still not going to lean on it: it is a company blog post with no byline and no methods section, credited to unnamed in-house data scientists with an Oxford researcher named as a collaborator, and it has never been peer reviewed; 273 and 79 do not add up to 360, and the average difference it reports growing to 38 days cannot be a phase difference at all, since a phase difference cannot exceed half a cycle. As far as we can find, there is no peer reviewed large scale analysis of app data on this question. Evidence that agrees with you is still evidence you have to check.

And the pheromone?

In 1998 Stern and McClintock reported in Nature that compounds collected from women's underarms could shift other women's cycles: shorter by 1.7 plus or minus 0.9 days with follicular phase compounds, longer by 1.4 plus or minus 0.5 days with ovulatory phase compounds, in twenty recipients from nine donors. They called it “definitive evidence of human pheromones.”

Strassmann's published reply is worth reading whatever you conclude. It notes that the result rests on a P value of about 0.05 with a sample that small, and that with twenty subjects one or two could carry the whole effect. It also asks, of the five cycles set aside because the women had mid cycle nasal congestion, what criteria were fixed in advance and whether the analysis was done blind. On that last point the 1998 paper deserves better than the insinuation, and gives its own answer: it reports the numbers with those cycles put back in, and the effect survives them, at minus 1.4 plus or minus 0.9 days and plus 1.4 plus or minus 0.5. The exclusion is a question about procedure, not the thing holding the result up. Reviewing the whole area in 2015 in Proceedings of the Royal Society B, Tristram Wyatt concluded that there is no bioassay led evidence that any of the four molecules routinely sold and cited as human pheromones is one, and ended: “It may be that we will find that there are no pheromones in humans. But we can be sure that we shall never find anything if we follow the current path.”

Bigger numbers than 1.7 days do appear in this literature, and they are the reason the third bench sits where it does. Russell and colleagues in 1980 reported a mean difference between a sweat donor’s onsets and their subjects’ falling from 9.3 days to 3.4; Preti and colleagues in 1986 reported the same quantity falling from 8.3 to 3.9. Those look like effects of five and four days, far larger than 1.7. They are not shifts in cycle length. They are reductions in the difference between menses onsets, which the 1986 paper calls in so many words “days’ difference in menses onset.” That is the same before and after gap the third bench runs on women who cannot see each other, and reports a convergence of three and a half days there with nothing at all to converge on. Spread over the four cycles of the Russell protocol, the implied per cycle movement is about a day and a half, below the figure the fourth bench is calibrated against. Whatever those experiments found, the statistic they found it with could not tell them.

Which is the fair statement of where this stands. Humans are mammals and may well have chemical signals; nobody has isolated one and shown what it does. But notice that the fourth bench never needed that question answered. Grant the signal, at the largest per cycle size we can find claimed for it, and the arithmetic still says it cannot hold two ordinary cycles together.

Then why does everyone remember it happening?

Because it does happen, constantly, and it needs no cause. Two independent clocks running at roughly the same rate pass close to one another all the time. The fifth bench counts what a year of that actually contains.

Bench five · a year in a shared flat

thirty pairs, twelve months
counted over 3000 unrelated pairs

Nearly every pair has their bleeding overlap at some point in a year. More than half get a month where the two of them start within a single day. And most of the same pairs also get a month where they are a fortnight apart, which is the identical evidence pointing the other way, and which nobody writes down, because there is no folk belief that periods repel.

That asymmetry is the whole engine. The coincidences are memorable and have a name. The non coincidences are unremarkable and have none, so they never enter the tally. It is not a failure of anybody's reasoning; it is what happens when only one of two outcomes has been given a word.

Worth adding, since this is a claim about what women believe: we could not find a well powered survey of that belief. The figures in circulation come from a study whose sample size we could not verify and from a set of twenty interviews. So the honest version is that the belief is clearly widespread and that nobody has measured how widespread.

What would actually settle it

Not a before and after gap. That statistic was asked to detect coupling and barely can, as the third bench shows in front of you. The right question is whether there is any coupling at all, and coupling has a signature that does not depend on where anyone started: a population of oscillators genuinely pulling on one another concentrates in phase, and the concentration is one number.

That is the second meter on bench three, and it is the same order parameter the corpus already uses on crowds of oscillators that really do synchronise. It responds to the dial. Anybody with onset dates from a group of women can compute it, and its null distribution can be simulated in a second on the machine you are reading this on.

Strassmann's Dogon analysis is the closest thing in the record to that test done properly, and it came out flat.

Where this page stops

Four things, because an argument is only as good as its edges.

This does not show that no coupling exists. It shows three narrower things: that the classic statistic barely distinguishes coupling from no coupling across the range of effects anyone has reported, that a pull the size of the largest published shift cannot hold typical pairs together against their own variability, and that the lived experience is fully accounted for with no coupling at all. A small effect, on some pairs, in some conditions, is not excluded by anything here, and Wyatt's review is right that the question of human chemical signalling is open.

The simulated women are a model. Each cycle length is drawn independently around its owner's average. That is an approximation: Ecochard and colleagues found in 2024 that cycle lengths do carry some dependence, though not at lag one, where the correlation was not significant. Cycle lengths are also right skewed rather than normal, and we draw from a normal truncated at 15 and 60 days. Real cycle length distributions carry a heavier right tail than that, so the model understates how irregular women actually are, which works against the argument here rather than for it. The truncation itself is doing almost nothing at these parameters: the bounds sit nearly four standard deviations out, and the simulated one month drift lands on the untruncated closed form to within a hundredth of a day.

The coupling model is the one most generous to synchrony. It is a phase response curve, the standard shape for biological entrainment, and it always pulls toward agreement and never away, on every signal, with no delay and no cost. A real signal would do less.

The parameters are published population figures, not ours. They are listed in the sources, and every one of them is on a slider, which is the point of the sliders. Where a choice could flatter the argument we took the other one: a within woman variability of 3.9 days from the Apple study rather than the 2.6 days from Natural Cycles, which is lower because that paper analysed only cycles in which ovulation was detected from daily basal body temperature, with LH tests as an optional extra, and that selection drops the irregular tail; and a bleeding duration of 4.0 days rather than the 6.2 days from diary studies, because the shorter figure makes the coincidence rates we report smaller. The offline verifier re runs the central result across the whole plausible box.

The argument is not ours. The ceiling and the forced return are Wilson's, from 1992, and Strassmann's. Simulation work on the same artifact was published by Jeffrey Schank between 2000 and 2006; those papers are paywalled and we could not read them, so we cannot say whether the particular quantities on this page appear in them, and we are not claiming priority over work we have not seen. What is offered here is a reproduction from scratch that you can operate.

The check

Recomputing…

quantityshipped in this pagerecomputed in your browser

Every uncertainty and free choice on this page.

Offline mirror, run before publication: node research/the-drift-that-has-nowhere-to-go/verify.mjs. It imports the same engine.js this page runs, re-derives every figure in the table above from fixed seeds, checks the closed forms against the simulation, sweeps the whole plausible parameter range, and reads this file to confirm that the numbers printed in the prose are the numbers the engine produces.

Sources, with identifiers