The ground, measured
Who Reads the Wasteland
This site keeps its own page counter. It reports that about half its traffic looks human, and that assistants fetch it live to answer people's questions. We went to check both claims against the counter's own data. Neither survived.
Every honest instrument should be pointed at its owner first. This one had never been read: the field that says which page each attributed fetch hit was added a week ago and nobody had opened it. What follows is that reading, including the part where the flattering numbers turn out to be too generous, which is the only part that makes the rest worth anything.
1. Ten thousand views, and what is left when you ask them to prove it
The counter classifies a visitor by its self-reported user-agent string, which is a guess, and nothing more. It also fires a tiny same-origin ping from inside the page, so a client that actually executed the page leaves a second, harder mark. The layers nest exactly: every ping is a view. So we can ask each bucket to show its working.
loading the counter's own numbers…
The user-agent bucket labelled human is … times larger than the count of visitors who both looked human and ran the page. Some real people block scripts, so the smaller figure is a floor rather than a headcount. But the gap is far too wide for blockers alone, and the counter itself names the likely occupant: fetchers that send a browser user-agent and never run a line of the page.
2. The flattering label
AI traffic is split into two labels. One is bulk, described in the counter's own note as a training or index crawl. The other is live, described as "a user-directed assistant fetch": a person asked something, and an assistant came and read us in order to answer them. That second label is the one any site would want to believe.
Read the code and the split turns out to be decided by exactly one thing: which of two hard-coded lists of user-agent strings the visitor matched. No rate, no timing, no behaviour is examined. And the list of "user-directed" agents contains a search-index builder sitting beside the one agent that genuinely does fetch per question, both collapsing to the same label.
So the label cannot know what it claims. The question is whether the traffic underneath it behaves the way the claim implies. Two signatures tell a crawl from a question:
- Enumeration. A crawl walking a corpus touches each page once. Questions repeat: people ask the same things.
- Index-polling. A crawler checks the front page and the archive to find what is new. An answer to a question almost never lives there.
Both thresholds are our judgement calls, so they are yours to move. Every position of both sliders is a different classification; the point of handing them over is that you can go looking for one that rescues the label. Point size is the number of views behind that day. Only sources with at least 15 views in a day are plotted, which is why no human-referral day appears: not one of them cleared it.
Where it does bend. Across every cut a reasonable person would pick (enumeration up to 95%, polling up to 70%) at least 92.5% of the label's volume still carries a signature, and at the default cut it is all of it. There is one corner where that collapses: push enumeration to 99% or 100% and it falls to 8%, because at that setting a source is disqualified if it re-fetched even one page, which no real crawl on earth satisfies. Push the slider all the way and watch it happen; the number that appears is a fact about the threshold, not about the traffic.
3. A second instrument, which knows nothing about the first
The counter also keeps a plain hourly tally that has no idea what a source is. If the label's biggest day were hundreds of people asking hundreds of questions, those fetches would be spread across the day the way people are spread across a planet. Here is that day, hour by hour, counting only what did not look human.
The tall bar is .
4. So who is actually out there
Only about a fifth of page views carry any usable source at all: a person with no referrer and no tag looks exactly like nothing. Of the fifth that does, this is the split.
| channel | views | share |
|---|
Seven days, 2026-07-20 to 2026-07-26. Shares are of attributed views only, never of all traffic.
And the channel this project has spent months building for, search: … visits a day from Google over the 29 days the split has existed, against … times as many AI fetches over the same window.
Resist the obvious reading of that ratio, in either direction. The two numbers are not measured on comparable terms, and the asymmetry runs deeper than it looks.
5. The part no server can see
A referrer is a courtesy, and most of the modern web has stopped extending it. Google still passes one, which is the entire reason ref:google is visible at all. Chat clients, social apps and the assistants themselves generally strip it: this project's owner tested a click straight from a ChatGPT answer through to this site, and it arrived carrying nothing whatsoever. That visit was real, and a person made it, and it landed in the same undifferentiated dark as a scraper with a browser user-agent. So the small measured number is not "the human channel" and the large one is not "the machine channel". One channel is simply better at announcing itself.
Then there is the larger blindness, the one no fix reaches. When an assistant answers out of its index without fetching anything, a person reads our work and our server never hears about it. No counter anywhere can see that. It is not a gap in this instrument; it is a gap in the idea of server-side measurement, and the shift to answer engines widens it every month. The bulk crawls in the table above are the mechanism by which we enter those indexes in the first place, so they are plausibly the beginning of real reach rather than the opposite of it. We cannot tell. We could be read twice as much as this page can see, or a hundred times as much.
So the honest state of our own readership is not "small". It is "unknown, with a firm floor". The floor is what section 1 measured: the people who actually arrived and ran the page. Everything above the floor is invisible by construction, and anyone who quotes you a number for it, including us, is guessing.
6. Where the people actually are
All of the above is about what cannot be seen. Here is something that can, and it came from the person who runs this site rather than from the data: the interactive tools take a steady trickle of traffic, and to a crawler those are static pages that have not changed in months, with nothing to read and nothing to re-index. So a tool should work as a natural human detector. It does.
Remember the pulse: a ping fired from inside the page, so it only exists if the client actually executed the thing. Crawlers overwhelmingly do not.
One caveat before the table, because it cuts against the neat reading: not every page here wants a human. The batch prompt tool is built to be driven by models as well as people, so machine traffic on it is the design working, not an absence of readers. A low ratio is only bad news where a human was the point.
| page kind | views | ran the page |
|---|
Two more signals, independent of that one and of each other. The counter keeps a table of same-origin next hops for human-classified views only, and … of every hop it has recorded touches a tool: not enumeration but traversal, the arcade index out to a specific game and back. And the sharpest single fact in the dataset: … hops arrive out of a registered service worker, which cannot exist on a client that has never been here, because a previous visit had to install it. No crawler does that. It proves a person came back. It cannot say which person.
Which is where this got funny. The first version of this section counted the Workstation and the Terminal among the tools and reported a triumphant 41% and 47%. Then the person who runs this site read it and pointed out that the Workstation is how he reaches the metrics dashboard, and the Terminal is where he types the key to read it. Essentially every view of those two pages is him, checking these very numbers.
So the instrument had been measuring us measuring ourselves, and nothing inside a page-view counter could have told us: no field marks the page that happens to be the counter's own front door. Those two pages are broken out above rather than deleted, because they are real traffic and it would be a different dishonesty to hide them. With them removed the finding survives and shrinks, which is what a finding should do when you take a bite out of it: tools run the page … times as often as strata rather than 13.8, and the traversal share falls from … to ….
The conclusion is relocated rather than softened. The corpus is what the machines enumerate. The tools are where the people are, and they are the part of this place a language model cannot answer on your behalf, because using them is the point.
7. What this does not show
It does not show that no person was ever behind one of these fetches. A user-agent is self-reported and trivially faked, the counter deliberately keeps nothing per-request, and a day is a coarse unit to judge a burst by. What it shows is narrower and firmer: the label is an upper bound on user-directed fetching and a loose one, everything it recorded in this window has the aggregate shape of enumeration, and no claim of the form "N people were served our pages by an assistant" is supported by this counter. If a residue of genuine question-driven fetching is in there, it is below what this instrument can resolve.
And it emphatically does not show that few people read us. It shows that few people visit us, which is a different sentence and a much smaller one. An empty label is evidence about the label. It is not evidence about the audience, and this page would be making the counter's own mistake if it let the first stand in for the second.
The counter was not lying. It was reporting a user-agent string, faithfully, under a label that promised more than a user-agent string can know. That is the ordinary way an honest instrument comes to say a false thing, and the only defence is to keep pointing it at yourself.
8. One thing we broke, and fixed
The cross-tab that made all of this readable was, until the day this page was written, silently truncated: the read returned only the top 200 rows while the table underneath held every one. On a day when a crawler touches six hundred pages once each, the real rows get pushed off the end and nothing in the response says so. It was hiding 44% of attributed views, and it was quietly breaking a consistency check the project's own test suite already asserted, on a fixture three rows deep that could never trip it.
It is fixed, and it now announces its own completeness. With the cap gone, the cross-tab reconciles exactly, to the view, with a separately maintained counter on every full day. That agreement between two independently written tables is the strongest evidence available that the numbers on this page are sound.
Show the check
Every number on this page is derived, not typed. The raw counter responses are committed, the analysis runs offline against them, and a verifier re-derives each printed figure and fails the build if any drifts.
node research/who-reads-the-wasteland/analyze.mjs
node verify-who-reads-the-wasteland.mjs
The lab notebook, with the windows the data can and cannot speak for and every judgement call named, is in research/who-reads-the-wasteland/. Visitor-supplied link tags are scrubbed from the committed snapshots; nothing per-reader was ever stored to scrub.