The Flood Nobody Measured

A river's flood peak is not weighed. It is read off a curve, from the height of the water. So we asked the United Kingdom's national flood record a plain question: for every published flood peak, had anyone ever actually measured a flow that big at that gauge? At 95.9% of stations the answer is no.

There is no scale under a river. To say that a flood carried so many cubic metres a second, somebody measures the height of the water, which is easy and continuous, and converts it with a rating curve, which is built by going out in a boat, or on a cableway, or wading in, and measuring the actual discharge at a known height. Those field measurements are called gaugings. The curve is only as good as the gaugings it was fitted to, and gaugings get made when a person can safely reach the river.

Floods are when they cannot.

Take the Aire at Armley, the gauge on the edge of Leeds. On 27 December 2015, during the Boxing Day floods, the archive records the peak at 344.4 m³/s at a stage of 5.217 m. That gauge holds 582 gaugings, from 1978 to 2025. The largest discharge anyone has ever measured there is 173.3 m³/s, on 2002-08-02, at a stage of 3.495 m. So the figure in the national record for that flood is 1.99× the largest flow ever measured at the gauge that produced it, at a water level 1.72 m above the highest one anyone has ever stood in the river to check. Ten more years of gaugings have been made since, and none of them has come near.

Pick a river

Every dot is one gauging station, placed at its own coordinates, which is why the shape of Britain appears: the country is drawn here entirely by the places somebody chose to measure a river. Colour is how far the biggest flood on record sits above the biggest measurement ever made at that gauge.

Station picker and stage-discharge plot

The plot is drawn from data your own browser has just fetched from the National River Flow Archive's public web service. Nothing on this page is our copy of it. If the archive is unreachable from where you are, the plot will say so.

What the record looks like from above

910stations with a published peak-flow record and at least one gauging
44,538annual maximum flood peaks in those records
35.6%of those peaks are larger than every gauging ever made at their own station
95.9%of stations where the biggest flood on record was never matched by a measurement

The median station's record flood is 1.93× the largest discharge anyone has measured there. At 287 stations even QMED, the median annual flood and the quantity the whole British flood-estimation method is anchored on, sits above every gauging in the archive.

Each station's record flood divided by its largest gauging, on a log scale. Left of the line, somebody measured a flow bigger than the biggest flood in the record. Right of it, nobody did.

The obvious objection, and what happens to it

Most stations were reading water levels long before anyone started gauging them: the annual maxima can begin in 1883 and the gaugings in 1988. So perhaps the count is measuring nothing but the late start of the measurement record. Restrict it to peaks that fall inside their own station's gauging span, where that objection cannot apply, and 12,628 of 36,712 peaks are still above every gauging: 34.4%, against 35.6% over the whole record. The finding survives its own most obvious confound almost unchanged.

Nobody has measured a flow that fills the channel

The archive publishes, for some stations, the bankfull flow: the discharge at which the river begins to leave its channel. That is a physical fact about the site rather than a property of how long the record happens to be, which makes it a cleaner question than any ratio. Of the 403 stations that have one, at 217, which is 53.8%, the largest discharge ever measured is below bankfull. At those gauges the whole out-of-bank range, which is to say every flood anybody worries about, has never been directly measured at all. The Aire at Armley is one of them: bankfull there is 232 m³/s and the largest gauging ever made is 173.3 m³/s.

By the size of the river

A single national fraction hides the thing that drives it. Large rivers get gauged into their floods; small ones do not, because a small catchment's flood is over in hours and nobody is standing there when it happens.

catchment, km²stations record flood above every gauging all peaks abovemedian ratio

Note the fourth column. The fraction of the record that is extrapolated falls by a factor of seven from the smallest catchments to the largest, and yet at every size of river, from moorland streams to the Thames, the single biggest number in the record was almost never matched by a measurement.

Somebody else already knew, and we can check against them

The National River Flow Archive grades every peak-flow station for what its records can honestly be used for: suitable for pooling, suitable for QMED only, or suitable for neither. The definition is a human judgement, in the archive's own words, that the station is acceptable if its largest annual maxima are “likely to be within 30% of its true value”. The assessors' reasoning is not published, only the verdict.

So here is a question that does not depend on us being right about anything: does a purely mechanical count, of how much of a station's record sits above its own measurements, line up with what the assessors decided?

the archive's gradestations median gaugings held record flood above every gauging all peaks abovemedian ratio

Monotone across all three grades, on every column that measures intensity. The stations the archive is happiest with have more gaugings, less of their record above them, and a smaller gap at the top. Two independent things agreeing is worth more than either alone: the assessors were tracking something real, and so is this count.

And look at the column that does not discriminate. The record flood sits above every gauging at around 95% of stations in every grade, the best included. Whatever the assessors are grading, it is not whether the biggest number in the record was ever measured. Almost nowhere was it.

By what kind of station it is

The comparison means different things at different gauges. At a velocity-area station the rating is nothing but a curve through the gaugings, so above them there is nothing else holding it up. At a weir or a flume the rating also rests on the structure's hydraulics, calibrated in a laboratory, and the gaugings are a check on that rather than its whole basis. So a weir extrapolating past its gaugings is standing on something; a natural section doing it is standing on less.

codewhat it isstations record flood above every gauging all peaks abovemedian ratio

The check that has to come first

A count like that is worthless unless the two things being compared belong in the same space. They might not: a gauging and a published peak could be on different datums, or the archive's stage series might not be the stage the rating reads. So before counting anything above the measurements, we checked the part below them, where both exist and can be laid against each other.

For every annual maximum whose stage falls inside the gauged range, we interpolated the station's own gaugings to that stage and compared the result with the flow the archive published for that event. If the rating is sound where it can be tested, the two should agree.

The check

Across 795 stations and 24,835 annual maxima that fall inside their station's gauged range, the published flow is a median 1.011 times what that station's own gaugings say at the same stage, and the typical station agrees to 5.4%. The gaugings themselves scatter about their own local line by 5.0%.

So the rating agrees with the measurements to roughly the precision of the measurements. Where it can be checked, it works. That is the whole reason the rest of this page is interesting rather than alarming: the instrument is good, and a third of the flood record lies outside the range where anyone has ever held it up against anything.

It is not uniform. At 612 of those stations the agreement is inside 10%. At 20 the published flows run more than half as large again as the gaugings at the same stage, which is a sign not that the archive is wrong but that the gaugings there disagree with each other. Station 25019, the Leven at Easby, is the clearest case: it holds gaugings of 5.5 m³/s at a stage of 0.504 m and 1.9 m³/s at 0.769 m, two measurements that cannot both describe the same rating. A measuring authority sees such points and sets them aside. Our arithmetic cannot, and says so.

Who already half-knew this

None of this is a discovery about hydrology. Extending a rating past its gaugings is normal, documented practice, and the people who do it say so plainly. The Environment Agency's own best-practice manual on the subject (Ramsbottom and Whitlow, 2003, W6-061/M) opens by naming the reason it is necessary:

“the extrapolation of rating curves to higher flows is hindered mainly by difficulties in actually measuring extreme flood flows to verify any such extensions. These problems relate to access, heath and safety and timing of measurements to coincide with flood peaks.”

and the National River Flow Archive's own accuracy page says:

“Uncertainties in the stage-discharge relationship can also be substantial, particularly in the extreme flow ranges where relatively few gaugings may be available to define the rating.”

What is odd is that three separate groups have counted a neighbouring version of this and each has filed it as an aside. Coxon and colleagues (2015), working on 500 UK stations with the Environment Agency's gauging archive, report that they “were unable to calculate uncertainties at low (high) flows for 31% (44%) of the groups of stage-discharge measurements”, because their method would not extrapolate past the measurements. That 44% is the closest existing UK number to the one on this page, and it arrives as the reason a method could not run rather than as a result. Mailhot and colleagues (2025), on the Quebec network, note that “the maximum gauged stages are smaller than the 5-year annual maximum measured water level for 52 % of the stations”, and add, in the same sentence, “(this analysis was not presented for conciseness)”. Gharari and colleagues (2024) count, month by month across Canada, the stations whose recorded stage sits above the highest gauged one, in one figure of a paper about something else.

So the honest description of this page is not that nobody noticed. It is that everybody noticed, in a parenthesis, and nobody counted. If you want to know what the extrapolation is worth once you are inside it, the two papers to read are Di Baldassarre and Montanari (2009), who measured mean errors of 1.2% interpolating against 11.5% extrapolating on one modelled reach of the Po, and Kiang and colleagues (2018), who put seven methods on the same data and got 95% uncertainty widths of 3 to 17% at median flows but 41 to 200% for high flows in an extrapolated section. Above the gaugings the experts do not merely become uncertain. They stop agreeing with each other.

What this is not saying

It is not saying the flood figures are wrong. Extrapolating a rating is a normal, considered, documented part of hydrometry, and the people who do it have tools this page does not use: the surveyed cross-section, the slope of the water surface, hydraulic models of the site, high-water marks left by the flood itself, and the behaviour of the structure the gauge sits on. A number can be well founded without ever having been measured directly.

What it is saying is narrower and, we think, worth being able to see: the published flood record and the measured record are two different things, and the gap between them is not small. On the evidence held in the national archive, most of what we know about how big British floods get is knowledge of a kind that has never been checked against a flow of that size at that place.

An honest catch, and why this page fetches its own data

The archive holds the gaugings its measuring authorities have supplied to it. It is not a register of every gauging ever made. So the correct reading of every number here is above every gauging in the archive's holdings, not never measured by anybody, and each count is an upper bound on the extrapolation rather than a certainty. What can be said about the size of those holdings is that they are not thin: the stations analysed here carry 242,970 gaugings between them, a median of 197 per station.

Two further limits worth stating rather than burying. The archive says that the ratings used to produce the peak-flow series “are not always the same as those applied in the main hydrometric archives” of the measuring authorities, so these gaugings are not guaranteed to be the calibration set for these peaks. And the comparison means different things at different gauges, which is why the table above is split by station type: at a velocity-area station the rating is nothing but a curve through the gaugings, while at an ultrasonic or index-velocity station discharge is not read off a stage-discharge curve at all.

A trap in the archive's own service, found on the way

The station metadata carries two fields naming water years and periods the archive has itself set aside as unfit for flood frequency analysis. The time-series service serves those years anyway, with nothing in the response to mark them. Measured here on 4 September 2026: of 2,211 rejected water years that fall strictly inside a station's served span, 2,047 come back in the data stream. So anyone who takes the peaks endpoint at face value is analysing events the archive has disowned. This census drops them, 2,536 events across 298 stations. Keeping them instead, the naive way, moves the headline from 35.6% to 35.4% and the station figure from 95.9% to 96.0%, so the choice is not load-bearing here. It is still worth the archive knowing, because for somebody else it will be.

The licence, and what it changed

The archive's API licence says at 3.1, quoted exactly:

“You may not make the Data available for download from any internet site without prior agreement..”

and at 3.2 bars distributing it to a third party “except as part of a product or application utilising the API”, while 5.1.1 requires an acknowledgement on “all copies of the Data, publications and reports” and 5.1.2 points at ordinary academic referencing. So the analysis may be published and the series may not be redistributed. That is why this page ships our arithmetic, the counts and the ratios, and asks your browser to get the archive's own numbers from the archive, live, when you pick a station. It is the shape the licence names, and it has the side effect of being more honest: you are not looking at our snapshot of somebody else's measurements.

The same licence carries a confidentiality clause at 4.1 which, read literally, would forbid publishing anything at all, and would contradict 5.1.1 and 5.1.2 in the same document. We have read it as boilerplate, as the presence of an acknowledgement clause and academic-citation clause requires, and have asked the archive to confirm. If they say otherwise, this page comes down.

Reproducing this

research/the-flood-nobody-measured/fetch.mjs pulls the four series for every station with a peak-flow record, one request at a time, and analyse.mjs computes every number on this page. verify-the-flood-nobody-measured.mjs re-derives the census from the analysed table, checks the arithmetic against worked cases computed by hand, and re-fetches a sample of stations live to confirm the archive still says what we read. The cache of raw series is deliberately not committed, for the licence reason above; the fetcher rebuilds it in about half an hour.

Acknowledgement: Data from the UK National River Flow Archive. Peak flow figures are the archive's live web service; its separately distributed WINFAP Peak Flow Dataset was at Version 15, released 27 August 2026, and the archive notes that the two may differ.