Artificial Wasteland · vexillology × colour science

Far Enough Away

A flag is a claim about identity made in cloth, and it has to survive the journey to your eye. Walk far enough from any two of them and they become the same flag. This page walks all 18,915 pairs of the world's national flags away from you, one step at a time, and reports where each pair gives up.

The measure is not pixel arithmetic. It is S-CIELAB, the 1996 model that blurs an image by the spatial sensitivity of human colour vision before comparing it, which matters because your eye is far blurrier for colour than for brightness: a red-and-green boundary dissolves at a distance where a light-and-dark one is still sharp. Every number here comes from that model, ported from the authors' own code and diffed against it.

Loading the measurement…

1 · The families

All 195 flags, with two dials. Distance sets how small the flag is in your field of view, which is what the model actually consumes. Tolerance sets how alike counts as alike. Flags within tolerance of each other are drawn together, and the grouping chains: if A is close to B and B to C, all three travel as one family.

-degrees wide
-away
-families
-in the largest
What the tolerance number means, and the one value that is not arbitrary

One value on that dial is not a matter of taste. 2.3 ΔE*ab is the just-noticeable colour difference reported by Mahy, Van Eycken and Oosterlinck (Color Research & Application 19(2), 1994, 105–121): below it, two flat colours are the same colour. Turn the dial all the way down to 2.3 and almost every family dissolves, because at that strictness only five pairs out of 18,915 ever merge, at any distance on the ladder. That is a result, not a broken instrument, and section 4 lists them.

Above 2.3 the dial stops being a threshold and becomes a question: how alike do two flags have to be before you would call them alike? There is no correct answer, which is why it is a dial and not a constant. Two honest stretches apply even at 2.3: that figure is a threshold for two large flat patches and this is an average over a whole image, and no human was asked. The ordering of pairs is far more trustworthy than any single line drawn across it.

The model works in angular size and knows nothing about metres. The distance readout is a convenience that depends entirely on how big the flag is, so it is a control rather than a constant.

The families, at this distance and this tolerance

- flags are no longer alone. Click any flag to load it into the comparator below. The on-screen shrinking is an illustration of the trend, not a calibrated stimulus: this page cannot know your screen or how far you sit from it. The comparator can, if you tell it.

2 · Two flags, at the size your eye would actually get

Pick any two. This draws them at their true angular size for your screen, so you can check the model against your own vision. When the readout says merged, look at them and see whether you agree.

-degrees wide
-away
-ΔE between them

The curve is that pair's whole life: how different the two flags are at every rung of the ladder, from a flag filling your view down to three arc minutes, where a flag is barely a shape at all. The dashed line is the merge threshold.

Calibrate this to your actual screen

By default the page assumes about 33 pixels to a degree, which is what a 96 dpi display gives at 50 cm. Your screen is probably not that and you are probably not sitting there. Hold a bank card against the screen and drag until the bar matches its long edge, which is 85.6 mm by international standard, then say how far away you are.

3 · Does the instrument work?

A model that ranks 18,915 pairs can always be made to look clever afterwards. So the target list was fixed first, from sources that exist for their own reasons: the Flags of the World reference page Country flags almost identical (last modified 2019), Britannica's Flags That Look Alike, and the pairs NAVA's own design booklet names as offenders. Together they name - pairs, out of -. That list was committed to this repository before the analysis was ever run against the corpus, and the commit order is the evidence.

-median rank of a named pair
-of them in the model's top 100
-AUC (0.5 would be nothing)
-worst-ranked named pair

A high-ranking pair that no source names is not a false positive. These lists are not exhaustive, so they bound recall and say nothing about precision. Pairs the model ranks high that nobody has written about are the interesting output, not the error, and they are listed separately below.

Every pre-registered pair, and where the model put it
RankPairMerges at (deg)

4 · The most confusable flags on earth

Ranked by how large a pair can be and still be merged. Flags marked not named by any source are ones the reference works do not mention, which is the part of this table that is actually new.

RankPairΔE at 2.2°Settles at

The five that actually become one image

Two flags can be alike in two quite different ways, and distance treats them oppositely. A pair can differ in drawing, like El Salvador and Honduras, two blue-white-blue tribands carrying different devices in the middle. Blur destroys drawing, so that pair converges all the way: past a certain distance they are not similar, they are the same picture. Or a pair can differ in shade, like Chad and Romania, whose layouts are identical and whose blues are not (#002664 against #002b7f). A spatial filter preserves the average of whatever it blurs, so a uniform difference in hue never washes out. That pair converges too, to 6.6, and stops there forever.

Which means the famous pairs are famous for the wrong reason. Of all 18,915 pairs, only these ever cross the just-noticeable line:

PairMerges atDistanceΔE up close

Three of the five are the Arab Liberation tricolours, which share a red-white-black horizontal triband and differ only in what sits in the white stripe: Egypt a gold Eagle of Saladin, Iraq the takbir in green script, Yemen nothing at all. The last row is a coincidence of averages rather than a resemblance: those two flags are 79.6 apart up close, and only meet at the very last rung, where each is a single dot roughly three arc minutes across. Ordinary acuity resolves about one arc minute, so nothing there is a flag any more. It is listed because excluding it would be a choice, and choices that tidy a result should be visible.

One more distinction the model cannot make and this page will not blur. What is measured here is simultaneous discriminability: two flags side by side, can you separate them. That is not the same as identification from memory, which is what actually happens when someone mistakes Chad's flag for Romania's, and which nobody wins, because nobody carries a calibrated blue chip in their head.

5 · Pairs separated by nothing but the shape of the cloth

The matrix resamples every flag onto one grid, so proportion is deliberately thrown away: the question is whether the design is confusable. Which means the model can then name the pairs whose designs are identical and whose only remaining difference is how long the flag is. On a pole, in wind, that is not much to go on.

PairProportionsΔE as designs

6 · The advice that fights itself

The standard guidance on flag design is Ted Kaye's "Good" Flag, "Bad" Flag for the North American Vexillological Association. Its five principles open with Keep It Simple ("so simple that a child can draw it from memory"), continue through Use 2–3 Basic Colors and No Lettering or Seals, and end with Be Distinctive or Be Related. Kaye states in Raven 8 (2001) that the principles "are generally non-overlapping as well as all-encompassing."

Three of the five can be measured off the flag files themselves, and the fifth is what this whole page computes. So the claim of non-overlap is testable. Here is each measure of simplicity against how far a flag sits from its nearest neighbour.

Simplicity measure vs. distance to the nearest other flag
MeasurePearson rSpearman ρ
Principle 3 directly: colours against distinctiveness
ColoursFlagsMedian ΔE to nearest

What the booklet does say, under principle 5: "Sometimes the good designs are already 'taken'." It concedes scarcity without ever connecting it to principle 1. The one place it comes close is a list of exceptions on page 14, where Maryland's "complicated heraldic quarters produce a memorable and distinctive flag" is offered as a departure from the rules rather than as evidence about them.

7 · Sixteen countries have two flags, and it changes the answer

Wikipedia's gallery lists sixteen states twice, once for the flag citizens fly and once for the flag the government flies, and the two usually differ by a coat of arms in the middle. There is no single right choice: nobody calls the German flag the one with the eagle, and nobody calls the Spanish flag the one without the arms. The rule used here is the page's own: take the row carrying the state's header cell, unless the other row is captioned "National flag of X", which happens once, for Spain. Both flags are in the corpus either way.

StateUsed hereAlso measured

Show the check

What counts as a flag here, and what was left out

The corpus is the 193 UN member states plus the two permanent observers, the Holy See and the State of Palestine, which is 195. That universe is not neutral and is not this page's invention: it is the one Wikipedia's Gallery of sovereign state flags uses, and the parse was cross-checked against the UN's own member list, with zero states unmatched in either direction. Ten entities the source lists separately under "Flags of de facto states" were dropped, across eleven rows because Transnistria carries two: Abkhazia, the Cook Islands, Kosovo, Niue, Northern Cyprus, the Sahrawi Arab Democratic Republic, Somaliland, South Ossetia, Taiwan and Transnistria. Each is recorded by name, with its reason, in manifest.json.

That boundary changes results, so it should be visible rather than assumed. Taiwan's flag appears in the vexillological source this page is validated against, grouped there with Samoa, and it could not be tested because it is not in this corpus. Wherever a pre-registered pair named something outside the 195, the pair was dropped rather than quietly counted, and the dropped names are listed in prereg.mjs.

The model is the published one, and that was proved rather than claimed

S-CIELAB is Zhang & Wandell's, and the authors released their MATLAB. That code is vendored unmodified in this repository and run under GNU Octave, so the JavaScript port can be diffed against it number for number rather than described as faithful.

CaseConditionsReference mean ΔEPort mean ΔEMax difference

The first row is the one that matters: where no deviation applies, the port and the original agree to the last bits of a double. The other rows are two deliberate departures, and the last row adjudicates them.

The reason is worth stating, because it is a real limitation of the published code rather than a quibble. Its padding routine refuses to pad by more than half the image, so once the eye's blur kernel is wider than the picture, the convolution runs off the end of its own padding and loses light. That regime is not an edge case here: the eye's filters span about a degree, and a flag seen from far enough away subtends less than that, so every interesting distance in this study is in it. The port lets the mirror reflection fold instead of truncating, which preserves a flat field exactly.

A shortcut that was tried, measured, and thrown away

The pairwise loop looked like it could skip most of its samples at the far end, since a blurred image is band-limited and Nyquist says sparse samples suffice. The argument is sound and the conclusion does not follow: the quantity being computed is an average, and Nyquist governs reconstruction, not the standard error of a mean over sixty points. Measuring it settled it.

Angular widthDecimationGrid leftWorst error (ΔE)

At the far end the error would have been several times the 2.3 threshold this page calls "merged". The shortcut is off; every pair is computed from all 24,576 samples.

Two errors this study made, and how they were caught

Seventy-two flags that were not flags. The first fetch wrote every HTTP response body to disk under a .svg name without checking what it was. Wikimedia rate-limited the run, and 72 of the 195 "flags" were 1,966-byte HTML error pages named chad.svg, brazil.svg and so on. Nothing errored and the directory listing said 195. An unchecked write makes failure indistinguishable from success.

Four flags that were the wrong flag. Andorra, Argentina, Haiti and Spain were taken from rows the page captions "Civil flag of…", which for those four means the version without the coat of arms. All four fetched cleanly, rendered cleanly and checksummed cleanly. Andorra's civil flag is a blue-yellow-red vertical tricolour with nothing on it, which is to say it is Chad's flag and Romania's. Had this gone unnoticed the headline here would have been a spectacular near-tie between four countries, and it would have been false. No checksum can tell you a correctly fetched file is the wrong picture. It was caught by making a contact sheet and looking at it.

Folklore this study checked and would not repeat

What this does not measure

Reproduce it

Everything is in research/flags-at-a-distance/: fetch-flags.mjs and repair-fetch.mjs build the corpus from Wikimedia Commons with a SHA-256 per file; scielab.mjs is the model; scielab.test.mjs is its self-check; cross-check.mjs is the diff against the original MATLAB; prereg.mjs is the answer key, committed first; analyse.mjs builds the matrix and report.mjs reads it.