EU DSA Transparency Database · seventeen platform-days · read 2026-08-30
Two Hundred and Forty Reasons a Second
European law makes an online platform file the reasons for its content moderation decisions into one public database. On 2026-08-30 the Commission's own front page reported 3,761,203,712 statements of reasons in the trailing 180 days, which is 241.8 a second. Less than two hours later the same block reported 3,774,070,771. This page takes whole platform-days out of that database and counts how many different things they say. Shopify's 199,335 decisions on 2026-08-28 carry one distinct public description between them. Pinterest's 596,796 carry 206, or 201 once case and spacing are folded together. Six of the seventeen days are recounted from the Commission's own zip file in your browser, by two engines that have to agree before a number is printed.
The filing, not the letter
Article 17 of the Digital Services Act says that when a platform takes your post down it owes you a reason. Article 24(5) says the platform must also send that reason to the European Commission, which publishes it. Both sentences are short, and the second one is the one this page is about.
Providers of online platforms shall, without undue delay, submit to the Commission the decisions and the statements of reasons referred to in Article 17(1) for the inclusion in a publicly accessible machine-readable database managed by the Commission. Providers of online platforms shall ensure that the information submitted does not contain personal data.
Regulation (EU) 2022/2065, Article 24(5)
That is the whole paragraph. It takes out personal data, and nothing else. So the thing in the database is not the notice a person received: it is that notice with the personal data removed. Whatever this page counts, it is counting the public record, and every sentence below is about the record rather than about what any user was told. The distinction is not a hedge: it is the difference between a claim this data can carry and one it cannot.
It is worth being exact about what the law does not say, because the obvious defence of a thin record is that the law made it thin. Article 17(3)(b) requires the statement of reasons to give the facts and circumstances relied on in taking the decision, and Article 17(4) requires it to be as precise and specific as reasonably possible under the given circumstances. The Commission's schema carries that through: decision_facts is a required free-text field with 5,000 characters to write in. Across the seventeen platform-days counted below, 2,225,655 statements, the longest decision_facts anybody filed is 638 characters, against a schema limit of 5,000; the longest typical value, the largest of the seventeen platforms' median lengths, is 191. The other four free-text columns have their own smaller allowances, and the longest value filed in any of the five is 1,431 characters of incompatible_content_explanation, which the schema caps at 2,000. So one filer did use most of one field. It is not a one-off in the sense you would want, though: that single longest account was filed 303 times, word for word, in a file that ships with this page so you can go and find it. The room is there. What the statute subtracts is names, not sentences.
The record has a shape. A submitted statement of reasons has 38 columns. Twenty of them are coded: the Commission's schema fixes each to one of a closed list, and there are 20 statement categories, 79 keyword values, 8 content types, 16 restriction codes and 2 decision grounds, a three-valued automation flag and a subset of 30 countries to choose from. Five of them are free text, and those are the ones Article 17(3) asks for in words. The remaining 13 are identifiers, timestamps, a reference URL, a product code and the _other escape hatches. None of the thirteen is an explanation, though two of them are the dates this page comes back to below. The largest of the five is decision_facts, which the Commission's own schema documentation calls a required textual field to describe the facts and circumstances relied on in taking the decision, and allows up to 5,000 characters.
This page calls those two halves the form and the prose, and counts them separately, because they turn out not to behave alike.
How many distinct sentences a day carries depends on which of the five columns you count, and that is this page's one free parameter, so here is the narrowest reading of it. Take only decision_facts, the single column the Regulation names in so many words. Pinterest's 596,796 decisions on 2026-08-28 carry exactly 200 distinct values of it, and the commonest one is filed 135,463 times, which is 22.7% of the platform's day in one sentence. Widen to all five free-text columns and it is 206. The counter below uses the five; the raw archives are here for anyone who wants a different reading.
One day, one platform
Pick a platform and a day. Everything in the panel is computed here, in your browser, from a table of class sizes: the number of statements is the sum of that table, and it has to come out equal to the count the Commission advertises next to the file on its own download page, or the row says so.
The counter
Loading.
The largest classes, with their own words
Recount it from the Commission's own file
Six of these days ship the original zip archive next to this page. Pressing the button fetches it, checks its SHA-1 against the checksum the Commission publishes beside it, unzips it here, and counts it again with two separately written engines. A number the two engines do not agree on is not printed.
What is inside the two hundred and forty
The rate in the title is real arithmetic on the Commission's own summary block: 3,761,203,712 statements over the trailing 180 days that block covers is 241.8 a second, averaged over six months, from 366 platforms. That block was read twice on 2026-08-30, at 07:17 and at 09:02 UTC, and in between it rose to 3,774,070,771, or 242.7 a second, a move of 12,867,059 in an hour and three quarters. That move is not a count of arrivals. The block is a trailing window, so what it reports is whatever came in minus whatever fell off the back of it, and on a different pair of days the same block can fall while statements are arriving. It is also behind the files it summarises: 3,761,203,712 over 180 days is 20,895,576 statements a day, which is fewer than the database took on any single one of the 49 days counted lower down this page. The figure moves while you look at it, it lags, and it is the least informative sentence anyone can say about this database. The reason is one line of the same source.
On 2026-08-28 the whole database took 27,104,899 statements. Google Shopping filed 17,045,654 of them, which is 62.89% of the day from one product-listing surface. The six largest filers are 90.48% of it. The listing carries a row for 366 platforms, and only 159 of them filed anything at all that day: the other 207 filed nothing. The 366 per-platform counts sum to 27,104,899 exactly, which is the number the Commission advertises for the global file, so the breakdown is complete rather than a sample.
And across the 49 days sampled here, the most recent page of the listing for each series with the one day Google Shopping has no file for removed, Google Shopping's daily count sits between 16,838,647 and 17,203,862: a spread of about 2% over seven weeks, while the database's daily total swings from 22,904,281 to 45,159,915. Its share of the day runs from 38.06% to 74.87% with a median of 59.86%, so 2026-08-28 is an ordinary day rather than a chosen one. A quantity that flat is a pipeline, and the sentence 240 moderation decisions a second is mostly a description of that pipeline.
Is that the form's doing?
Here is the objection, and it is the right one:
You have counted a dropdown menu. The database is a compliance form with a closed vocabulary and a legal ban on putting personal data in it, so of course a day of it holds only a couple of hundred different answers. The number is a property of the form, not of anyone filing into it.
That objection makes a prediction, and the prediction is testable on the same day's files. If the format is what sets the ceiling, then the ceiling is the same for everyone: the free-text half should be the flattened half for every filer, because that is the half Article 24(5) bites on, and the amount the prose adds beyond the coded fields should be near zero everywhere.
The quantity to look at is exactly that: how many bits the free text adds once you already know the coded fields. Written out, it is the conditional entropy H(prose given form), and it equals H(form and prose together) minus H(form), both computed from the class sizes in front of you. Zero means the prose is a function of the dropdowns: knowing the coded values, you already know the sentence. The counter above reads it out directly, in its own box, for whichever platform-day is selected, and it is the last column of this table.
| platform | day | statements | form | prose | both | impossible dates | bits: form | bits: prose adds |
|---|
The prediction fails, and it fails in both directions on one day. Facebook, Shopify and Fortnite add 0.000 bits: for those three, the free text is a strict function of the coded fields and carries nothing the dropdowns did not already say. eBay adds 2.055, Pinterest 1.931, LinkedIn 1.875. Nor is it that the free text is always the poorer half: Snapchat's twenty coded columns separate its day into 178 classes while its five free-text columns separate it into 17, and eBay's do the opposite, 87 against 279. Same schema, same statute, same 24 hours. A ceiling imposed by the form cannot be in two places at once.
Facebook is the case where the mechanism is visible in the text itself. Its 803,293 statements that day carry 7 distinct free-text descriptions across 209 distinct coded ones, and not one coded class has more than one sentence attached to it. The largest of the seven, filed 366,522 times, reads in part:
This was DECISION_ACCOUNT_SUSPENDED because it violated DECISION_GROUND_INCOMPATIBLE_CONTENT.
incompatible_content_explanation, Facebook, 2026-08-28, file 224237
The field Article 17(3)(e) asks for an explanation in has the schema constant from the field next to it pasted into a sentence. That is why the number is 0.000 and not merely small, and it is a filing choice, not a constraint: the same schema on the same day let eBay put 2.055 bits there.
A larger number in that column is not a compliment. It means a bigger lookup table, not bespoke prose. eBay's 279 free-text classes cover 24,259 statements, about 87 statements per sentence, and its longest decision_facts value that day is 366 characters against the 5,000 the schema allows. Nobody in this file is writing to anybody. The question the column answers is narrower and more useful: how much of what the record could carry did each filer put in it.
Two more columns of the same record
The same seventeen files answer two smaller questions that need no interpretation.
The schema has two date fields. content_date is defined by the Commission as the upload or posting date of the content; application_date is the date that this decision starts from. A decision cannot start before the thing it is about was posted. On 2026-08-28, all 28,653 of GitHub's statements say the content was posted on 2026-08-28 and the decision started on 2026-08-26, the same pair of dates on every row of the day. Dailymotion has 165 of 5,836; Snapchat has 1 of 13,171; the other 14 platform-days have none. So it is not a property of the database. It is one filer's clock, and the count is the impossible dates column of the table above.
The other is smaller and more literal. On 2026-08-28, 118 of Reddit's 1,434 statements, and 178 of the 1,402 it filed the day before, contain an unsubstituted template placeholder in the middle of the sentence: A third party submitted a {ip_content} takedown notice that affected the following. The variable was never filled in. It is a small thing, and it is the whole argument in one string: what reaches this database is a template with slots, and here is one that nobody filled.
What the count is, and what it is not
Two hundred and six sentences for six hundred thousand decisions is not, on its face, a breach of anything. Article 17(3) asks for the facts and circumstances, the ground, the automation disclosure and the route to redress, and the Commission's schema delivers most of that through coded fields by design. Nothing in the Regulation requires a different sentence per decision, and as precise and specific as reasonably possible under the given circumstances is a standard no row count can adjudicate: for a day of near-identical product listings, one sentence may be exactly as specific as the circumstances allow. A page that read 206 as non-compliance would be wrong.
What the count does establish is narrower and, in a database this large, more useful: it is the resolution of the public record. On 2026-08-28, six hundred thousand Pinterest decisions arrive in the record as 362 distinguishable things, and the average one of them is indistinguishable, on that record, from 52,116 others. Nineteen bits of address would be needed to point at one of them; the record carries 4.31. That is the ceiling on every question anyone will ever ask this database, however many rows it grows: no analysis of the public copy can separate two decisions that arrived in it as the same row of text, and for one platform, on one ordinary Friday, that is 199,335 decisions arriving as one.
The check
Two published anchors per file, neither chosen by this page. Every
one of the seventeen archives was re-derived from its raw bytes on 2026-08-30, and for
every one the recount equals the statement count the Commission advertises beside it
on its own download page, and the SHA-1 of the bytes equals the checksum the
Commission publishes beside it. Seventeen for seventeen, on both, and those two
anchors are re-checked against the manifest on every run whether or not the bytes
are present. What a run in a fresh checkout cannot do is re-derive the
counts for the eleven archives that do not ship: the Commission publishes a
checksum and a row count beside each file but no distinct-value count, so for those
eleven the numbers in the table rest on the derivation until you download the
archives and pass --raw. That is the weakest joint on this page and it
is why the raw files for six of the days are here. The
366 per-platform counts on the listing also sum to the
Commission's separately published global total, exactly, on both days read. When you
press recount above, your browser does the hash check itself and prints
it.
An independent route that disagrees, and the disagreement is the finding. The distinct-prose count was taken twice: once on the exact strings, and once after NFC normalisation, whitespace collapsing and case folding. On fifteen of the seventeen days the two agree exactly. On the two Pinterest days they do not: 206 against 201, and 235 against 228. Every one of those differences is a single letter of case in the phrase Advertising Guidelines against Advertising guidelines in the contractual-ground field: the verifier folds those two spellings together and requires every remaining pair to become identical, so a second kind of difference hiding inside that gap would turn it red. The exact count runs 5 ahead of the meaning-preserving one on the first day and 7 on the second, and the page prints both rather than picking the flattering one.
Two engines in your browser, and a refusal. The recount runs two
separately written implementations of one written specification over the same bytes,
one parsing the CSV with a character state machine and counting into hash maps, the
other splitting records by quote parity and counting by sorting and run-length. The
panel is painted by /_kit/concur.js, whose guard() returns
nothing when the two engines part, so a contested number never reaches this page.
Agreement here is not proof: two implementations can share one wrong reading of a
specification, and nothing in this panel can see that.
A third implementation, in another language. The verifier re-runs the whole protocol in Python, using Python's own zip and CSV libraries rather than the hand-written zip reader and CSV scanner the Node side uses, and compares every field. The protocol is frozen in a comment at the top of both files, written before any result was read.
The control on the control. The verifier plants 50 synthetic rows carrying explanation strings that occur nowhere in the real file, re-runs the same unmodified counting function, and requires the distinct count to rise by exactly 50 and the statement count to rise by exactly 50; then it restores the file and requires the original numbers back. It also corrupts a published SHA-1 and requires the hash check to fail, deletes ten rows and requires the advertised-count anchor to fail, and rewrites a figure in a copy of this page's own HTML and requires the display check to go red. A checker that cannot go red is not a check.
And a control on the statistic that carries the argument. The column above is the whole second layer, so it is attacked directly rather than around the edges. Shuffle the free-text columns between rows, which destroys the coupling between the two halves and changes nothing else, and the bits added must rise by more than one; replace every row's free text with a strict function of one coded column and they must fall to exactly 0.000, with no coded class left holding two sentences. A number that were hardcoded, or that quietly ignored one of its two halves, would sail through the anchors above and die on this pair.
Every figure printed above is string-matched, and the instrument is
driven. The verifier strips the scripts, styles and comments out of this
file and sweeps every run of digits left in the readable text, on this page
and on the sharing card and the structured-data block that travel without it. Each
one must either be a figure this run recomputed, matched inside a phrase from its own
sentence, or sit on a short published list of things that are not measurements at
all: dates, article numbers, file identifiers, pixel widths. The match runs in both
directions, so a figure in the prose that no computation produced is a failure, and
so is a computed figure that the prose quietly dropped. An earlier version of this
gate looked only at the figures it had itself marked up, which meant a number written
into a sentence without the marker was checked by nothing; that hole is what the
sweep closes.
Then it stands up a stub browser, runs this page's own controller against the data
this page ships, changes the platform and presses the keys, and compares what lands
in the readouts. That covers the numbers only a running page produces. Twenty
deliberate corruptions of the finished files were fed to it, one at a time, from a
digit changed in a sentence to a byte flipped inside a Commission archive to the
reintroduction of the engine defect it had already found: it caught
20 of 20, and was required to go
green again between every one. Five of the twenty exist because an audit of this page
proved the corresponding check could not fail, and one of those five caught a
replacement check that tested only half of what it claimed to. The harness is
research/two-hundred-and-forty-reasons-a-second/mutation-battery.sh,
and it works on a scratch copy, so anyone can run it.
The free choices, named. Which columns count as form and
which as prose is this page's decision, made once and listed in full in
research/two-hundred-and-forty-reasons-a-second/derive.mjs; the split
follows the Commission's own documentation, which divides the attributes into
free textual and limited, the value provided needs to be one of the
allowed options. The normalisation is NFC, whitespace collapsing and case
folding, and nothing else. The twins figure is the expected size of a
record's own class, which grows with the size of the day and is not comparable
between platforms of different size; the bit figures are, and both are shown.
Where the comparison would be vacuous. On a platform-day with one class, every count is 1 and every entropy is 0.000, and the conditional-entropy column cannot distinguish the prose adds nothing from there is nothing to add to. Shopify on 2026-08-28 is exactly that case, and it is left in the table with its 1 rather than dropped, because it is the strongest single observation on the page and the reader should be able to see that it is also the degenerate one. Fortnite is the other end of the same trap: one prose class over 15 coded ones, where 0.000 is forced. Facebook is the case that carries the argument, because it is the one that is not degenerate: 7 free-text classes spread over 209 coded ones, no coded class holding two of them, and 0.000 bits added over 803,293 statements. GitHub, SoundCloud, Discord and the two Reddit days are the other non-degenerate near-zero rows, at 0.003, 0.008, 0.018, 0.011 and 0.044 bits.
Run it.
node research/two-hundred-and-forty-reasons-a-second/verify-two-hundred-and-forty-reasons-a-second.mjs.
Six of the seventeen archives ship with this page and are re-derived from raw bytes
every run; the other eleven are re-derived if you download them and add
--raw <dir>, and the verifier prints how many of the seventeen it
re-derived and how many it skipped rather than passing quietly on six.
The battery of deliberate corruptions is a second command,
bash research/two-hundred-and-forty-reasons-a-second/mutation-battery.sh.