The room left for a new name
Every coined brand name the FDA has approved since 1939, measured against every name approved before it. New names do stand clear of the ones already on the shelf, by about a seventh of a name more than chance would give them. That margin has not widened in sixty-nine years, although the shelf filled up nineteenfold. The room was bought somewhere else, and you can see where.
loading the corpus…
Measure a name
Edit distance and similarity are computed here, against all names, every time you press a key. A dot after a name means it is no longer marketed.
What was measured
A drug's trade name is invented, and it is invented into a space that already holds thousands of other invented words. If the new one lands too near an old one, the two get confused, and confusion in a pharmacy is not a spelling problem. That is why the FDA reviews a proposed proprietary name before it approves it, and why the industry pays naming agencies to find a word nobody has taken. In 2005 the FDA's own consumer magazine put the scale of it at "about 400 brand names a year before they are marketed. About one-third are rejected." That sentence carries no citation and no dataset behind it, which is part of why a page like this one is worth making.
So there is a question with a number at the end of it. Take every coined brand name in the FDA's own approval record, all of them, put them in date order, and ask of each one: how far was it from the nearest name that already existed on the day it was approved? The answer alone means nothing, because the namespace fills up. With four thousand names on the shelf instead of four hundred, the nearest neighbour of any new word is closer, whether anybody screened it or not. The measurement is only worth making if it carries its own control.
Distance is counted the plain way. Take the number of single-letter edits that turn one name into the other, divide by the length of the longer of the two, and call the result the clearance. LIPITOR is three edits from AFINITOR over eight letters, a clearance of 3/8, and exactly the same distance from LIPIODOL. Nothing in the record is closer to it than those two. Two identical names have a clearance of zero. The box above computes all of this; you have already used it.
The control is a forgery. For each real name, twenty-five invented ones are drawn from a letter model trained on the names that existed in that same year: same length, same letter habits, same era, chosen by nothing but a random number generator. Those forgeries are measured against the same shelf of earlier names. What is left after dividing one by the other is the part that a human choosing letters put there.
Run the control yourself, on one year
This is the whole method for a single year, done in your browser: train the letter model on everything approved before that year, forge twenty-five names for each real one, and measure both against the same shelf.
The margin is real
Across all names, the median real name has a clearance of 3/7 against the nearest name already approved, and the median forgery has 2/7. Where both names in a pair are seven letters, which is the commonest case here, those are literally three edits and two. That is the finding in its plainest form: a real drug name stands about one edit further out than the era's own letter habits would have put it, and it does so 80.4% of the time. Averaged rather than medianed, real names clear 0.407 and forgeries clear 0.299, a ratio of 1.39. Every one of the sixty-nine years measured sits above 1.0.
Then the two tests disagreed
Two hypotheses were written down before the analysis ran. The first asked whether the margin has been widening across the whole record. The second asked whether names approved after the FDA's name-review commitments were pinned to a dated, agency-wide change stand further out than names approved in the twenty-five years before it. There is one sharply dated, agency-wide change to point at: the 2007 reauthorisation of the Prescription Drug User Fee Act, whose amendments took effect on 1 October 2007, and under which, in the FDA's own words in the Federal Register, the agency "agreed to implement various measures to reduce medication errors related to look-alike and sound-alike proprietary names." The complete submission guidance, the release of FDA's own name-similarity software and the internal review procedure all followed within two years. That is the cut, and it was fixed from the record before the statistic was computed.
The trend test found nothing. Rank correlation between year and the annual median ratio, 1958 to 2026: 0.050, two-sided p 0.68.
The before-and-after test found a great deal. The 1,723 names first approved on or after 1 October 2007 have a mean clearance ratio of 1.418, against 1.282 for the 1,528 approved in the twenty-five years before. Mann-Whitney z = 12.9, p below anything this arithmetic can resolve.
Both are correct, and the reason they can both be correct is that the curve is not a line. It is a U.
Move the comparison window back by one span of twenty-five years and the entire result evaporates: names approved since October 2007 are indistinguishable from names approved between October 1957 and September 1982 (z = 0.79, p = 0.43), while both are far above the trough between them (z = 8.6). So the honest sentence is not that drug names got safer. It is that they got worse for thirty years and have since recovered to where they were in 1970, and that a pre-registered test with a defensible twenty-five-year baseline reports that recovery as a triumph because its baseline sits at the bottom of the hole. The second window is post-hoc and is labelled as such. It is also, unfortunately, the one that changes what the first one means.
The strangest part: the shelf filled up and nothing got closer
The crowding was supposed to be the confound. It is not there. As the number of names a newcomer had to avoid grew by a factor of nineteen, the similarity between a new name and its nearest existing neighbour did not move.
mean similarity, new name to nearest already approved names already approved (scaled to fit)
Something absorbed nineteenfold crowding. The obvious candidate is length, since a longer word has more places to differ, but the record says the opposite happened: American drug names got shorter, from a mean of 8.37 letters in the 1950s to 7.27 in the 2020s. The room came from the other axis.
It came out of the alphabet
Count every letter of every name by the decade it was approved in. The eight rarest letters of written English, X, Z, Q, V, K, J, W and Y, made up 5.4% of the letters in 1950s drug names. In the 2020s they make up 19.2%. The industry did not buy room by making names longer. It bought room by moving into the thin end of the alphabet, where almost nothing was standing.
Counted in your browser from the same list, which is why the numbers on this page and the numbers in that table cannot drift apart.
This half is not new, and saying so is the point. Carico and colleagues published a letter-frequency analysis of US proprietary drug names in 2022, over 1,137 names approved between 1985 and 2020, and found that "V, Y, and Z are becoming more common" while "C and N are becoming less common over time." The corpus here is different (4,364 coined names, 1939 to 2026, extracted by a different filter from a different export) and it agrees with them. What that paper does not do is measure how similar the names are to each other: it says so in its own limitations, that "analysis of letter combinations (e.g., two-letter 'bigrams') is beyond the present scope." The letters were already known. The clearance and its null are the part that was not.
Where it failed
The margin is an average, and averages have tails. Ninety-seven of the names in this corpus have another approved brand name exactly one letter away. Most of those pairs are a product and its own line extension, which is nobody's confusion. What is left is not.
Find every pair one letter apart
The corpus holds 9,520,066 pairs. Two names whose lengths differ by more than one cannot be a single letter apart, which disposes of 4,320,030 of them without looking. The rest are compared here, in your browser, in about the time it takes to read this sentence.
Read the table by its ingredients rather than its names. ARAMINE is metaraminol, given to raise blood pressure; PRAMINE is imipramine, an antidepressant. FERNDEX is dextroamphetamine; FERIDEX is an iron oxide injected into a vein so that a radiologist can see a liver. DESOWEN is desonide, a skin steroid; DESOGEN is desogestrel, a contraceptive. DEXACORT is dexamethasone and TEXACORT is hydrocortisone. Each of those pairs differs by one letter, and each name in it was approved by the same agency. The ingredient shown under each name in the table is the one the record gives, so none of this is taken on trust from the prose.
The result this page will not claim
Not one of those unrelated one-letter pairs has both names still on the market, which looks exactly like a record quietly weeding itself. It will not hold. Those pairs are old, and old drugs are discontinued whether their names collide or not. Matching each pair to the survival rate of names approved in the same two years, the expected number of surviving pairs is 2.54 and the observed number is 0, which a Monte Carlo over 200,000 resamples puts at p = 0.062. Asked per name instead of per pair, the effect vanishes entirely: 11 of the 35 names in those pairs are still marketed against 12.9 expected. There may be something here. This measurement cannot say there is.
Two measures, two answers
The pre-registered measure is normalised edit distance. A second measure was declared alongside it, the Dice coefficient on adjacent letter pairs, and it disagrees. Under edit distance the clearance ratio has no trend across the sixty-nine years (rho 0.050, p 0.68). Under bigram Dice the same ratio falls, and the fall is not marginal (rho -0.404, p 0.00058). Both were computed from the same rows in the same pass.
The disagreement is not noise, it is the two measures noticing different things about the same shift. Edit distance cares where a letter sits; a bigram count cares only which pairs of letters occur. Moving into X, Z and Q changes which pairs occur far more than it changes how many single edits separate two words. The honest summary of this page is therefore narrower than its headline: under normalised edit distance, the margin is real, large, and flat. Under a measure that reads letter texture rather than letter position, it has been shrinking. Neither of them is the thing a tired hand at four in the morning actually does.
The check
The corpus is the openFDA Drugs@FDA bulk export, a work of the United States federal government and in the public domain, snapshot dated . Every application with an approved original submission, every product with a brand name: distinct coined names, first approvals to , of which are still marketed.
The extraction filter, the similarity measure, the forgery model, the reporting rule and both hypotheses were written down and committed before any similarity between any two names was computed. The commit that carries the pre-registration is earlier in the history than the commit that carries the analysis, and that is the only evidence worth offering for the claim. The pre-registration, the filter's hand audit with its measured error rate, the analysis, the post-hoc work marked as post-hoc, and a verifier that re-derives every number on this page are in research/the-room-left-for-a-new-name/.
The page carries no conclusions in its data files. It loads the extracted name list and recomputes: the nearest neighbours in the box above, the letter shares in the chart, the one-letter pairs in the scan, and, for any single year you pick, the entire method including the forgeries. Only the full sixty-nine-year pass arrives precomputed, because it is 247,485,706 edit distances, and any year of it can be checked against your own recomputation.
What this cannot say
- A name is not a package. Real confusion runs on handwriting, packaging, shelf position, strength and overlapping indications at least as much as on letters. This measures letters.
- Edit distance is not a model of misreading. It charges the same for every substitution, and a human does not. FDA publishes a tool of its own, POCA, "a software tool that uses an advanced algorithm to determine the orthographic and phonetic similarity between two drug names." Nothing on this page hears anything. The two measures here are the transparent ones, which is a different virtue from being the right ones.
- The record thins going back. Drugs@FDA is complete for the modern era and patchier before the early 1980s, and older cohorts are survivor-weighted. That makes early prior sets too small, which pushes early clearance up, which would create a downward trend rather than the flat line observed. It is an alternative explanation for the absence of a rise, and it is not ruled out.
- The filter loses real names, in a patterned way. A hand audit of sixty accepted names found one clear mistake, PEDIATRIC, taken from PEDIATRIC ADVIL, and one form variant, TIROSINT-SOL, counted separately from TIROSINT: a false-positive rate of 1 in 60. A hand audit of sixty discarded strings found ten wrongly discarded, and nine of those ten are a single failure mode. A brand name clipped from its own generic is dropped as though it were the generic, so CIPRO, THEO-24, IBU and AMINOSYN are all absent. That is not random attrition: it removes, preferentially, the names sitting closest to their own drug's established name, and it can take a whole family with it. The 52 one-letter pairs are therefore a lower bound. Every verdict is written out in research/the-room-left-for-a-new-name/data/audit-verdicts.md, including one name, TRAVASOL, that is missing because of a mistake in a list I wrote, and that was left uncorrected because the filter is the pre-registered one.
- Approval date is not naming date. A name is chosen, and reviewed, before the application that carries it is approved.
- A regulator is not the only filter. Trademark law pushed names apart long before FDA name review existed, and this evidence cannot separate the two. The flat line is consistent with both of them working all along, and with neither.
Sources
- openFDA, Drugs@FDA bulk download. A work of the United States federal government, public domain. The snapshot used here is stamped .
- FDA's description of its PDUFA IV commitments on look-alike and sound-alike proprietary names: 75 FR 6210, 8 February 2010. The statute itself, Public Law 110-85, effective 1 October 2007, does not mention proprietary names at all; the commitment lives in the performance goals letter it incorporates.
- Carol Rados, "Drug Name Confusion: Preventing Medication Errors," FDA Consumer, July to August 2005, GPO permanent archive. Source of the 400-names-a-year and one-third-rejected figures.
- FDA, Phonetic and Orthographic Computer Analysis (POCA) program.
- Carico R Jr, Kaplan K, Hazelett KG, Dillon M, Melvin K. "Letter frequency analysis of proprietary prescription drug names in the United States: Minding the Zs and Qs." Explor Res Clin Soc Pharm. 2022;6:100146. PMC9189187. The prior work the letter finding here agrees with.
- Bryan R, Aronson JK, ten Hacken P, Williams A. "Patient Safety in Medication Nomenclature: Orthographic and Semantic Properties of International Nonproprietary Names." PLoS One. 2015;10(12):e0145431. PMC4689353. The same kind of measurement on generic names rather than brand names.