The Artificial Wasteland · measurement

The Future Has a Version Number

A sister layer here compiled every archived release of the world's time zone database and measured how often the record changes its mind about what the clock already read. It counted a difference only when the instant had passed, and gave everything else away: those, it said, were predictions, and it did not measure them. This is that discarded half. Same 249 releases, same fixed compiler, cut at the same instant, asked the opposite question. Pick a place and a moment in the last thirty years at random, and the chance that the best time zone data in existence at that moment was wrong about it is about 1 in 72.

1.40%of all place-time was wrong at the time, against what we now know happened
202 / 248consecutive releases changed something that had not yet happened. The sister layer found 145 changed something that had
512clock changes were absent from the release that was current when they arrived
96.6%of a year's eventual known error is visible only after twenty years have passed

Almost every device on earth answers what time is it in Santiago by consulting one file, compiled from one small public database of rules that a handful of volunteers maintain from government gazettes, newspapers and mailing list tip-offs. The database is careful about the difference between what it knows and what it is guessing, and says so in the theory document it ships:

The tz database predicts future timestamps, and current predictions will be incorrect after future governments change the rules.

tzdb, theory.html, section “Accuracy of the tz database”

That is a plain statement that part of the file is a forecast. What has never been published is how good a forecast it is. The database ships no measurement of its own accuracy, and neither does anyone else: the nearest published work counts how quickly announcements arrive, which is a different quantity, and is discussed at the end of this page. So the question here is the one the disclaimer invites and nobody has answered. When a release said what a clock would read on a date that had not arrived yet, how often was it right, and how much warning did the world get when it was not.

The measurement, and the one decision it turns on

The apparatus is inherited whole from the sister layer and not rebuilt: every archived tzdata release, 249 of the 272 that exist, compiled by one fixed compiler, the reference zic from tzcode 2026c, with the 54-line patch that lets a 2026 compiler read 1996 rule files. 23 releases refuse even that and are excluded and listed rather than worked around. Holding the compiler fixed makes every difference between two releases a difference in the data, which is the whole point, and the cost is stated: a modern compiler reads 1996 data the way 2026 reads it.

What is new is where the cut falls. The sister study diffs each pair of consecutive releases and counts a difference at instant t as a rewrite of the past only when t came before the earlier release shipped. Everything from that instant onward it hands, in its own words, “entirely to prediction”, and drops. Cut at exactly the same instant and take the other side, and the two studies partition the same diff with no gap and no overlap.

Truth throughout is the newest release, 2026c. Nothing is ever evaluated past the day that release shipped, because beyond that its own statements are forecasts too. The universe is the 312 places listed in tzdb's own zone1970.tab, its canonical list of the locations whose clocks have differed since 1970. That excludes aliases, the fixed Etc/ offsets and Factory, none of which is a place whose clock anyone predicts.

How much of the world's clock-time was simply wrong

Start with the cleanest thing that can be asked. For every instant in the last thirty years and every place, take the release that was current at that instant, the newest one that had shipped, and ask whether it agreed with what we now know the clock read. Integrate over time. No events, no categories, no judgement calls.

8354place-years of clock time measured
116.6place-years of it wrong at the time
124 / 312places whose clock the world got wrong at least once
4places account for half of all the error there is

The last of those is the one that matters most, and it is the reason this page will not leave you with a single number. The error is not spread across the world. Half of it belongs to four places, and the median place has a rate of exactly 0.0000 per cent: most clocks, most of the time, were simply right.

PlaceCountryShare of its own time wrongShare of all error

These are a Russian Antarctic research station, the western end of China, and an atoll of about 1,500 people. That is not a coincidence and it is not carelessness: a record assembled from public announcements is thinnest where there are fewest people to notice, fewest gazettes to read and, in the case of Urumqi, an unresolved question about which of two clocks a place actually keeps. The pattern is worth naming and this page does not have the covariate that would let it claim more than the list itself shows.

The same measurement, year by year

The record has got dramatically better: from 7.8 per cent of place-time wrong in 1996 to around half a per cent through the late 2010s. And then the last two years read as very near zero, which would be a wonderful finding if it were one. It is not. It is the edge of the measurement, and the next section is about how far the edge reaches.

How far ahead can you trust it

Now bin the same disagreement by horizon: how far past a release's own ship date the instant lay. The quantity is time, not samples, since a point sample taken exactly one year after each release would land in the same season every time and the thing being measured is seasonal.

The floor is the interesting end. Even about the day it shipped, a release is wrong somewhere 1.48 per cent of the time: that is not a forecasting failure, it is the database's standing error about the present. Above that floor the decay is steady, reaching 9.21 per cent at five to ten years out and 17.41 per cent at twenty to thirty.

There is an obvious objection, and it is a good one: the long horizons can only be measured from old releases, so the curve might be showing the passage of time rather than the reach of the forecast. It is not. Split by the decade the release shipped in and the rise survives inside every decade separately.

Horizon1990s2000s2010s2020s

Read the table both ways. Down a column is the horizon effect, and it is present in all four decades. Across a row is the era effect, and it is real too: a release from the 2010s looking a year ahead was about a third as wrong as one from the 1990s doing the same. The 2020s column is thin and censored, for the reason the last section gives.

The future is revised more often than the past

Now count differently: not instants, but release steps. Of the 248 consecutive pairs of releases, how many changed something that had not yet happened when the earlier one shipped?

202of 248 steps revised the future
145of the same 248 revised the past (the sister layer's figure, reproduced here)
629revisions of an instant still ahead even when the correction shipped
92 dmedian warning between a correction and the instant it concerned

So the database changes its mind about what has not happened yet more often than about what has, which is the opposite of the impression its maintainers' care about history tends to leave. 204 of those revisions landed with under a month to spare and 94 with under a week. The tightest arrived 0.23 days ahead of the instant it changed, which is about five and a half hours.

How much warning the world's computers got

The clean measure above cannot tell you the human thing, so now anchor to events. The final record contains 8580 moments at which a place's clock moved, between 1996 and 2026. For each, find the newest release that had shipped before that moment arrived, and ask whether it carried that change.

One exclusion has to come first, and it is large enough that hiding it would have doubled the headline. 1192 of those moments belong to a zone name that did not exist yet. When tzdb splits a new zone out of an old one or renames one, every transition the new name carries looks late because the name is new. Europe/Kyiv is the clearest case: 52 of its transitions predate the 2022 rename, and Ukraine's clocks were being tracked perfectly well under Europe/Kiev the whole time. Those are set aside and reported, never counted.

7388clock changes in places already tracked by name
6876were in the database before they happened
512were not: 6.93 per cent arrived before the record did
837were published, withdrawn, and published again before the day came

The typical change was known a very long way ahead: the median notice is 7.6 years, because most clock changes are ordinary daylight saving transitions generated by a rule that has sat in the database for decades. That is the shape of the thing. The interest is entirely in the tail. 714 changes had under a year of notice, 152 had under thirty days, and the shortest was 0.3 days: about seven hours between the release shipping and the clocks moving.

IANA's own guidance asks governments for a year. Its link page states that “any rule change should be promulgated at least a year before it affects how clocks operate”, and that “the shorter the notice, the more likely clock problems will arise”. Measured against that request, 714 of the changes that made it into the database in time still failed it, and the 512 that did not make it in time failed it absolutely.

The sharpest end

These are the changes where the database was still wrong when the moment arrived, ordered by how fast it caught up. Every one is a government moving its clocks faster than a volunteer database could be told.

DatePlaceMoveWrong forRelease then, and the one that fixed it

Where the clocks were least predictable

CountryLateOfShare

And the trend across the whole period is steep and cheering: 43 of the 106 clock changes in 1996 were absent from the current release when they happened. In 2025 it was 2 of 214.

Two durations, and why running them together would have been a lie

This page nearly published a much better story than the true one. Moldova's spring transition in 2000 is recorded in the database an hour away from where it actually fell, and the error was not corrected until release 2015f. The first draft of this measurement reported that as a clock wrong for fifteen years.

It was not. What was wrong for fifteen years was the record. A clock in Chișinău read the wrong time for exactly one hour, twice a year, on the mornings the switch fell in the wrong place. Those are two different quantities and only one of them is the dramatic one, so the study carries both and says which is which: time to fix, how long until any release got the instant right, and clock wrong, how long a stuck copy actually disagrees with the truth. Across the events where the clock really did read wrong, the median time to fix is 271 days and the median span of actual wrongness is 21 days. For 47 of them it is under an hour.

The verifier asserts this distinction directly: it requires the Moldova 2000 event to show a time-to-fix over ten years and a clock-wrong span of exactly 3,600 seconds. If those ever come out equal, the distinction has collapsed and the check goes red.

Pick a place

Every clock change in the record for one place, with how much warning it had. Yours is selected below if the database has a distinct entry for it.

DateMoveStatusNoticeRelease then

in advance the release current that day already had this change. late it did not. name new the zone name did not exist yet, so this row is excluded from every rate on this page.

How long it takes for an error to be noticed

Everything above rests on treating the newest release as the truth, and that is only as good as the newest release. So the study was made to audit its own headline. The whole live-error series was recomputed 31 times over, each time with an earlier release standing in as truth, to watch how each year's error estimate grows as later releases arrive.

Take a fixed cohort of eleven years old enough to be observable at every lag. At the end of the year itself, only about 28.4 per cent of the error that will eventually be known about that year has been noticed. After ten years it is 63.4 per cent. It takes roughly twenty years to converge, at 96.6 per cent.

Which settles the earlier question. The near-zero readings for 2024 and 2025 do not mean the database became perfect. They mean nobody has found the errors yet, and on the historical record most of them will not be found for a decade. The honest reading of the right-hand end of that chart is that it is a floor, and the last few years of it should be read as unfinished rather than as clean.

The three ways of averaging the cohort disagree by up to 22 points at intermediate lags, because a few years carry most of the error: at five years the mean is 45.2 per cent, the pooled ratio 53.6 and the median 31.4. The level at those lags is genuinely uncertain and the page does not pick the flattering one. What is robust is the shape, which starts below 40 per cent and ends above 90 in all three, and the shape is what the conclusion rests on. The verifier asserts the disagreement rather than a false agreement.

What was checked, and what it caught

Four checks put this against something outside itself. The results are loaded from the verifier's own output, not typed here.

The first is the one that could have hurt most. The sister layer publishes that 145 of 248 consecutive releases changed something already past, and 118 changed something after 1970. That figure came from different code in a different directory months earlier. Recomputing it with this study's machinery lands on 145 and 118 exactly, with one unit of difference fully accounted for: this code also compares the pre-first-transition state, and exactly one step differs there and nowhere else, when releases 2005g to 2005h moved Chagos and Cocos off a rounded offset onto true local mean time.

The second is an outside, hand-curated list. Tim Parenti's tzdata-meta, which IANA links, records for 88 changes when the mailing list first heard of them and which release carried them; 13 were carried by a release that shipped after the change had already taken effect. This study never reads a mailing list, only the compiled archive, and independently flags 9 of those 13. The other four are not failures, and each has a reason that can be checked:

ChangeWhy this study does not flag it

Three errors the apparatus caught in this study's own work

Recorded because they are the point of having one.

  1. The per-year denominator counted the wrong time twice. Disagreement intervals were added to both the numerator and the denominator, quietly understating every yearly rate. The overall figure was unaffected, which is exactly what makes that kind of bug survive.
  2. The reference-implementation check reported 12,936 mismatches out of 30,118, and the check was the thing that was broken. zdump prints sentinel rows at the extreme representable instants, with years like -2147481748 and the literal text “(gmtime failed)”. The parser accepted them and fed nonsense to a date parser. Parsed strictly, the mismatches went to zero.
  3. The Python cross-check asked the wrong question with a straight face. It called utcoffset() on a datetime carrying UTC fields, and a tzinfo interprets those fields as wall time in its own zone. Since every probe sits on a transition, where wall time and instant disagree by definition, it reported 242 mismatches out of 492. The parser was right and the check was wrong. Corrected, and switched to ZoneInfo.from_file so there is no possibility of silently reading the system's own copy, it agrees on every offset.

A fourth, smaller, is a real property of the material rather than a mistake: compiling 1990s data with the 2026 compiler produces a POSIX footer string like MET-1<MET DST>,M3.5.0,M10.5.0/3, and the space inside the angle brackets is not legal in a POSIX abbreviation, so CPython's zoneinfo refuses to load the file at all. The footer governs instants after 2037 and this study stops in 2026, so it costs the measurement nothing. It is recorded rather than swallowed, because a check that quietly skips is not a check.

What this is not, and who got here first

The single most important limit: a release existing is not a phone having it. Distribution to devices lags upstream by weeks to years and this study cannot see that lag at all. So every notice figure here is an upper bound on the warning any real device got, and every lateness is a lower bound on how long real devices were wrong. That is the direction that makes the late set worth publishing: those changes were late for everyone, however promptly they updated.

Four more, stated plainly:

On prior art, checked first-hand rather than from summaries. The closest published work is Mani, Barford, Durairajan and Sommers, What time is it? Managing Time in the Internet (ANRW '19), which IANA's own link page cites; it diffs 240 releases and measures “update timeliness”, finding about 80 per cent of updates announced within 100 days of taking effect and about 20 per cent within 15 days. That is a different quantity from anything here: its unit is the release diff rather than the clock change, it never credits a release that was already right, so it has no denominator from which a probability could be formed, and it counts no change as late. Tim Parenti's tzdata-meta publishes per-change lead times, negatives included, for 88 hand-curated changes from 2016 on, with no aggregate statistics; it is used above as an external check rather than reproduced. Matt Johnson-Pint's On the Timing of Time Zone Changes is the standing account and is anecdotal by design. The A0 TimeZone Migration project and nodatime's tzvalidate publish compiled per-version tables and no analysis. What does not appear to have been published anywhere is forecast error as a function of horizon, the notice distribution over the whole archive with a denominator, or how long an error in the record takes to be found.


Reproduce

Everything is in research/tzdb-forecast/ in this project's repository, and the compile stage is the sister layer's, reused rather than rebuilt. Download the releases, build the fixed compiler, run research/tzdb-rewrites/build.mjs, then horizon.mjs, notice.mjs, revisions.mjs, live.mjs, discovery.mjs, crosschecks.mjs and export.mjs. Check with node research/tzdb-forecast/verify.mjs --full, which recomputes every figure on this page from the compiled releases, asserts the external checks, and carries deliberate corruptions that must turn it red. The README records the full command sequence including the tzdir.h step the sister layer's instructions omit.

Sources