Ground truth · the declared wall, dated
When the Wall Went Up
Ninety hosts this corpus cites now refuse to let it re-read them. The Internet Archive keeps their robots.txt like any other page, so each door can be dated, and the date comes out as a bracket rather than a day: after the last capture that still said yes, no later than the first that said no. The width of that bracket is the Archive's visiting rate, not the web's rate of change, and it is published beside every date.
This corpus cites 2,870 sources on 859 hosts. On 12 September 2026 a frozen parser asked each host's robots.txt whether this project may still read what it cited, and 90 hosts said no. That is a fact about one day. It says nothing about when the answer changed, and a series that starts in September 2026 can never find out, because it has no yesterday.
The Internet Archive keeps robots.txt like any other page. So the question has an answer, and the answer has a shape: not a date, a bracket. After the last capture that still said yes, no later than the first that said no. How wide that bracket is depends on how often the Archive happened to visit, which makes the measurement partly about the record rather than about the web. That correction is not ours. It was made in public by @melodic.stream on 13 September 2026, against a looser claim of ours, and it is the estimator this page registers and reports.
What was read
| Hosts asked | 180 | 90 that refuse us today, and a seeded 90 that do not |
| Hosts with any archived robots.txt | 158 | of 180 indexed |
| Monthly probes ruled | 8,700 | 321 unusable, carried as holes rather than dropped |
| Read to the same depth in both arms | 5,096 / 3,604 | probes in the refusing arm and the permitting arm. Two arms read to different depths are not comparing what they claim to |
| Captures that answered with a status other than 200 | 1,671,201 | counted separately, because a 404 is not an absence |
| Revisit records | 6,198,450 | the Archive noting it served bytes it already held; not a fetch and not a failure |
The front, host by month
Each row is one host, each column one month from 2021-01 to 2026-09. Warm is refused, cool is permitted, blank is a month the Archive did not visit. Rows are sorted by when the refusal begins, so the wall is the diagonal. Click a row to read the two files that changed the answer.
The matrix needs JavaScript. Everything it draws is in payload.json, and every number on this page is in the prose beside it.
refused permitted unreadable capture no capture names an AI crawler
Every date is a smear, and this is how wide
Of the 90 hosts that refuse us today, 55 have a two-sided bracket in the record: a capture that permitted and a later one that refused. 19 were already refusing at the first capture inside the window, so their wall predates what can be seen here and is reported as left-censored rather than dated. 13 refuse us live and never refuse in the archived record at all, and 3 have no usable capture.
The median bracket is 4.1 days wide. A tenth are narrower than 0.1 days and a tenth wider than 69.1; the widest is 1388.3. 46 of 55 come in under thirty days. That distribution is the Archive's visiting rate, not the web's rate of change, and it is the reason a date here is written as an interval and not as a day.
When the names arrived
The other half of the question is not about us. A host can refuse a crawler it has never heard of by refusing everyone, and it can also write a specific name into the file. Over 180 hosts, 73 have ever named one of the 23 tokens this study watches for. Two of those tokens are decoys, carried to see whether the wave is real.
First appearance, by token
| token | hosts | earliest here | median | its own docs first archived |
|---|---|---|---|---|
| Amazonbot | 46 | 2021-06-03 | 2025-06 | |
| CCBot | 51 | 2022-02-04 | 2024-10 | |
| magpie-crawler | 10 | 2022-02-04 | 2024-05 | |
| ChatGPT-User | 39 | 2023-04-01 | 2024-05 | |
| GPTBot | 66 | 2023-08-24 | 2024-01 | 2023-08-07 |
| Google-Extended | 45 | 2023-10-03 | 2024-06 | 2023-09-28 |
| Bytespider | 49 | 2023-10-10 | 2025-06 | |
| anthropic-ai | 36 | 2023-11-02 | 2024-11 | 2023-12-13 |
| cohere-ai | 29 | 2023-11-02 | 2024-11 | |
| Omgilibot | 24 | 2024-01-01 | 2024-10 | |
| PerplexityBot | 33 | 2024-03-01 | 2024-12 | 2024-01-09 |
| ClaudeBot | 51 | 2024-03-29 | 2025-04 | 2024-04-11 |
| ImagesiftBot | 21 | 2024-03-29 | 2024-11 | |
| Claude-Web | 26 | 2024-04-04 | 2024-11 | |
| Diffbot | 28 | 2024-04-16 | 2024-12 | |
| YouBot | 27 | 2024-05-29 | 2025-06 | |
| Applebot-Extended | 30 | 2024-07-04 | 2025-04 | 2024-05-08 |
| Meta-ExternalAgent | 37 | 2024-08-01 | 2025-07 | |
| Timpibot | 16 | 2024-08-01 | 2025-06 | |
| OAI-SearchBot | 16 | 2024-09-01 | 2025-04 | |
| AI2Bot | 15 | 2024-10-22 | 2025-07 | |
| Scrapy decoy | 14 | 2022-02-04 | 2025-02 | |
| FacebookBot decoy | 30 | 2024-01-01 | 2024-10 |
Sorted by the earliest date anything here names them, decoys last. The last column is a control, not a citation: the earliest archived copy of the operator's own page about that crawler. A robots.txt naming a token before its documentation existed would be a fault in this instrument.
Not everything in that table is an AI crawler, and the ones that are not are the reason to look twice. 5 of the tokens are already in somebody's robots.txt before OpenAI's own GPTBot page was first archived on 2023-08-07: Amazonbot (2021-06-03), CCBot (2022-02-04), magpie-crawler (2022-02-04), ChatGPT-User (2023-04-01), Scrapy (2022-02-04, a decoy). Common Crawl's CCBot has been in block lists since long before any of this, and Amazon's crawler is not an AI crawler at all. Averaging those into a wave would be doing the thing this page is about, so they are named instead.
A block list arrives as a list
The pre-registration asked which of two Anthropic tokens a host writes down first, expecting an order. The answer kept coming back neither, because they land in the same capture, and that is a better question than the one asked, so it was measured and is marked unregistered on the page as well as in the output.
Of the 68 hosts here that name two or more crawlers, 31 (45.6%) name every one of them in a single capture, and the median host's largest single-capture group is 8 tokens. The median host arrives at its list in 2 events, not one per crawler.
That is a fact about people rather than about crawlers. Nobody sits down and considers Bytespider on its own merits. A block list circulates, somebody pastes it, and a dozen refusals appear between one archive visit and the next. It is also why the dates on this page cluster: they are not a dozen independent decisions, and a count of refusing hosts is closer to a count of times a list was copied.
The instrument, checked against itself
Five controls were registered before any of this was fetched, and all five run on every pass. They are the reason a number here is worth anything, so they are on the page rather than in a notebook.
Controls
| A token cannot predate its own documentation | 6 | tokens with a dated operator page to check against; the earliest GPTBot capture anywhere here is 2023-08-24, and platform.openai.com/docs/gptbot was first archived 2023-08-07 |
| The id_ modifier returns the archived file | pass | one capture fetched both ways every run; the plain form must be markup, the id_ form must not, and they must differ |
| A zero must be able to be a one | 40.6% | share of hosts where the token probe found something, so a host reported as naming nothing is a reading and not a broken instrument |
| Hosts read at every distinct capture | 17 | the seeded control arm, to test what bisection assumes. 16 of 17 land on exactly the same capture as the monthly grid plus bisection. The 1 that do not are captures the Archive indexes and cannot replay: the dense arm needed a body that returns nothing, so what that disagreement measures is this harvest's failure rate, not the assumption. |
| Captures the Archive indexes but will not serve | 321 | of 8,700 probes (3.7%). Every one was asked for twice. They are carried as holes, which widens a bracket rather than narrowing it |
| The decoys | FacebookBot 30 · Scrapy 14 | two tokens carried that are not AI crawlers; if the block-list era is real they should not arrive with the wave |
The prediction that died most usefully
Prediction 8 said the decoys would not arrive with the wave: two tokens were carried that are not AI crawlers, Scrapy and FacebookBot, and their first appearances were supposed to be more scattered in time than GPTBot's. They are not. GPTBot's first appearances span 436.7 days between the first and third quartile; the decoys span 315.8 and 253.3.
That is the most useful dead prediction here, because of how it died. FacebookBot is not an AI crawler and it turns up in the same captures as the ones that are, which is exactly what the co-arrival result above says should happen: the boilerplate block lists that circulate carry it along with the rest. So the clustering this page measures is not evidence that publishers were reasoning about AI crawlers one at a time. It is evidence that they were copying a file. The registered control and the unregistered finding are the same fact seen from two sides, and the control is the side that keeps the finding honest.
What was registered, and what it did
Eight predictions were committed before the harvester existed, and the verifier checks that the pre-registration's commit is the older one. 4 of 8 scored predictions held.
Pre-registered predictions
| # | prediction | value | outcome |
|---|---|---|---|
| 1 | median first-refusal month for named-group CLOSED hosts falls in calendar 2024 | 2024-11 | held |
| 2 | more than half of two-sided CLOSED brackets are under 60 days wide | 89.1 | held |
| 3 | at least 20 OPEN hosts name an AI token in their most recent capture | 9 | KILLED |
| 4 | anthropic-ai first appears before ClaudeBot at more CLOSED hosts than the reverse | {"aiFirst":3,"claudeFirst":2,"sameCapture":25} | held |
| 5 | token sets are non-decreasing at more than 80% of hosts that ever name one | 76.7 | KILLED |
| 6 | at least 5 CLOSED hosts flap: refused, permitted, refused | 6 | held |
| 7 | at more than 10 CLOSED hosts the refusal predates any Claude token | 10 | KILLED |
| 8 | the decoys do not wave: their first-appearance IQR exceeds GPTBot's | {"gptIqr":436.7,"decoyIqrs":[315.8,253.3]} | KILLED |
What this cannot say
- It dates the declared wall, never the enacted one. What a robots.txt says is not what a server does. Measuring the second means presenting as a crawler we are not, and this project will not do that.
- Absence of a capture is not absence of a rule. A host the Archive rarely visited gets a wide bracket, and a host it never visited gets none at all.
- A host is not a publisher. Several hosts here belong to one company. Nothing on this page counts organisations.
- It cannot say why. A Disallow is a fact about a file, never a motive.