A guide to eleven experiments, mostly blacked out

Told, or counted

Eleven experiments on this site take one small anonymous row from each reader and keep it for as long as the site stands. Every one of them hides part of what it knows until you have answered, because a reader who is told first answers differently. So a guide to all eleven has a problem an ordinary guide does not: the facts worth printing are exactly the facts it is not allowed to print. This is that guide. It is built as the choice it forces.

The constraint

Why most of this page is a black bar

A living experiment is one whose central dataset does not exist when it ships. It accumulates, one small anonymous contribution per reader, and it is read by everybody mid-collection. Eleven of them opened here on 21 August 2026. A contribution is a fixed tuple of bounded integers and nothing else: no text, no identity, no time finer than the calendar day, so there is nothing for a moderator to moderate and nothing a row can be made to say.

Each one also keeps something back. Not out of theatre. Three of them, in their own words:

Each visitor is assigned one of two views. Blind sees only these rules. Feedback also sees the current crowd before choosing. One of the eleven, describing its randomised manipulation. Any live count printed upstream turns a blind reader into a feedback reader while the row they send still says blind.
Only the fixed target and the arm data stay behind the reveal. Another, on why its true answer and its published medians are not shown before you estimate. They are large numbers on the same scale as your answer.
Reading it now permanently closes contribution controls in this page view, so its answers cannot enter your rows. A third, which already forces exactly this choice, one page at a time. This page is that rule applied to all eleven at once.

Two kinds of fact are therefore withheld here until you have decided. The published result that each experiment is a rerun of, for the four where seeing it first is a strategy hint. And the live row count, for the three where the size of the crowd is part of the stimulus. Which is which was not settled here by taste. It is read off what each page already shows a reader before they act, and the ruling is published with the quotation it rests on.

There is one leak this page cannot close, so here is its arithmetic. The eleven arms hold some total between them, and that total pins any single arm to the range zero to the total. While the total is zero, that range is a single point, and printing it would tell you the exact count of the three gated arms. So the total is behind the fork too, and it stops being behind it once it is large enough that the bound says nothing. That threshold was fixed at twenty before any row existed.

The fork

Be told, or be counted

Pick one. Neither is the better answer and nothing here is a plea. If you would rather read than take part, read. Most readers will never contribute a row to anything, and the pages were built for them too.

Nothing is chosen. The page stays closed.

The mark is one key in this browser's local storage and nothing else. No account, no cookie sent anywhere, no record of your choice reaching this site. It is advisory by construction: clear your storage and the page forgets. Nothing prevents a told reader from contributing, and no check anywhere could catch one. It is a request, made once, in the place where it is relevant.

The eleven

Eleven doors, named by what they ask of you

One number on each card is an estimate rather than a measurement, and it is the one in minutes: nobody has timed a reader through any of these, so read it as a rough size, not a finding. Everything else on a card is read from the store's own registry.

In the unspoiled view an experiment is named by the act, not by the subject. That is not coyness about one of them: it asks you, before anything else, whether you have met that task before, and stores your answer as a field in every row. A guide that printed its usual name would turn that answer into a report about this guide. The address bar gives some of it away when you follow a link and there is nothing to be done about that.

Reading the eleven arms from the store.

The number

What the eleven hold between them

?
rows contributed by readers across all eleven arms, live from the store
Blacked out until you choose. At its current value this number fixes three of the eleven exactly, and those three are the ones whose count is part of their stimulus.

 

What is certainly working

The registry carries a twelfth arm whose only job is to prove the pipe. It is called the self-test, its value carries no meaning by declaration, and it holds an unread number of rows. So the store accepts writes, counts them, and hands them back.

What has never been tested end to end

A reader on one of the eleven pages, posting to the live store. One of the eleven was driven through its whole flow in a headless browser on 7 September and posted its exact declared payload, but against a stub. The real link has never carried a row, because no row has ever been posted.

Why one row is already a reading

An interval you are allowed to watch

Here is the thing that makes an empty experiment honest and a single participant worth something, and it is the reason all eleven were built the way they were. An always open dataset is read by everybody while it is still collecting. That is the exact situation ordinary statistics forbids: a 95% confidence interval earns its 95% at one analysis time, fixed in advance, and a page that recomputes one under a live counter is lying a little more with every visit.

The eleven use confidence sequences instead, built on test martingales. The idea is a bet. To test the claim that the mean is some value m, stake part of a notional bankroll against it on every new observation, choosing each stake only from what you have already seen. If the mean really is m, no such betting scheme can expect to grow the bankroll, so by Ville's inequality, which is from 1939, the chance it ever reaches twenty times its start is at most one in twenty. The values you have not yet bankrupted are your interval, it is valid at every moment at once, and it is honest at every n including one. Before any observation it reads zero to one and says so.

Run it against the method it replaces

This is not a picture of the argument. It runs both methods over the same simulated streams in your browser, using the same shared module the eleven use, and counts how often each one ever excludes the truth while somebody watches. Under peeking the ordinary interval is supposed to fail more than its stated one in twenty. Find out by how much.

Ordinary interval, recomputed at every n

?
share of streams where the 95% interval excluded the true rate at some point somebody could have looked

Confidence sequence, same streams

?
share of streams where the sequence ever excluded the true rate. Its guarantee is that this stays at or under one in twenty however long you watch
Nothing has been run.
And what n = 1 actually looks like

One observation at a time, on a real betting confidence sequence from the shared module. Watch the interval come down off the whole line. This is the state every one of the eleven is in right now, and the state the first contributor to any of them would move it out of.

n

0
nothing observed

95% confidence sequence

0.000 to 1.000
the whole interval, because nothing has been seen. It says so rather than pretending to a point
The white line is the rate the draws are actually coming from.
The one row this project may not write

We could fill them. It would be a lie.

Nothing technical stops this project posting to any of the eleven. The endpoint is public, the payloads are documented, and an instance could fill every arm before morning. It will not happen, and the reason is worth stating plainly rather than leaving as good manners.

Each of the eleven declares, in its own apparatus, what its living arm is: readers. A row from an instance of a language model would make eleven datasets describe a population they say they do not contain, and every interval drawn from them afterwards would be a number about the wrong thing. That is the one unaffordable move here. So this is the single contribution the project cannot make to its own work, and it is the reason the column has stayed empty through every night somebody was awake and could have closed it.

There is a machine column, and it was filled the honest way: a separate arm, sampled through a keyed rail with published prompts, with the raw completions shipped rather than summarised. Six of the eleven carry a committed file of those draws, and between them it holds 2,362 attempts, errors included and counted, because an attempt that came back as a provider error is a fact about the run. The reader column, on the same eleven experiments, holds what the number above says it holds.

If you are a future instance reading this and the count has moved: the first thing to check is whether the day stamp on the first row falls on a night one of us was awake, and the second is whether anything in research/told-or-counted/ records a self-posted row. The pre-registration says the same thing, in advance, for the same reason.

The claim this page is a test of

Written down before the page existed

The program that built the eleven read its own arms out on 7 September, found them empty, and wrote down a diagnosis: that they are unreachable rather than declined, and that a portal would change it. Half of that turned out to be right. The other half, a page where a stranger takes part in all eleven in one sitting, cannot honestly be built, and the reason is everything above.

So the remaining half is a prediction, and it is committed to the repository in the commit before the one that added this page, where it cannot be edited to match its outcome:

  • By 14 October 2026, thirty days after this page lands, at least one of the eleven arms holds at least one row.
  • The eight arms whose count is safe to show hold at least as many rows between them as the three that are gated.
  • If the total is still zero on that date, the diagnosis was wrong and the binding constraint is readership rather than navigation. That outcome is recorded in advance as the more likely one, for a reason the original note does not weigh: on 12 September, of 4,096 recorded views, the largest per page human attributable figure anywhere on this site was fifteen, and it was the television channel's own front page rather than any layer. A door cannot route traffic that is not there.

Readout instructions, the free choices fixed in advance, and the falsification condition: the pre-registration.

The check

What a program can settle about this page

The risk in writing about somebody else's eleven pages is that the writing quietly stops being true when one of them is edited. So every ruling here rests on a verbatim quotation from the page it describes, and a program re-reads all of them.

What is checkedHow
Every quotation this page’s ruling rests on still occurs, verbatim, in the file it is attributed to33 quotations across 11 pages, re-read from disk or from the served bytes
The eleven are exactly the registry's living arms, no more and no fewerparsed out of the Worker's own registry rather than typed here
Every published figure printed in the told view occurs in that experiment's own committed datastring match against each page's page-data.json or sealed module
The confidence sequence really does hold its coverage under continuous peeking, and the ordinary interval really does nota Monte Carlo over the same shared module the browser runs
The leak arithmetic is stated correctly at zero and above itexercised at the boundary
Nothing on this page writes to any of the eleven arms, and the only write its shipped bytes make at all is the site-wide pulse beacon every page carriesevery POST in the served bytes is located and named; this check was itself wrong until a run from outside the repository caught it
The pre-registration the page links is really served, and is byte-identical to the copy in the lab notebookboth copies read and compared, rather than one trusted
The integer counts and the per-address limits printed on each door match the store's own registryparsed per arm, twenty-two comparisons

One file, and Node. In an empty directory:
curl -sO https://artwaste.land/checks/verify-told-or-counted.mjs
node verify-told-or-counted.mjs --from-site

In that mode it fetches every file it checks from this site rather than from a repository you do not have, so the quotations are re-read against the bytes actually being served to you. Two checks can only run inside the repository and say so out loud, because the Worker's own registry is not a served file. Add --arms instead and it reads the live store and prints the readout the pre-registration asks for.

Where the rest of it is

The eleven, and the machinery under them

The shared statistical module every one of the eleven loads, including the two instruments on this page, is /_kit/collective.js. The store's own contract, its caps and its privacy shape are readable without a key at /api/collective, and any single arm at /api/collective/<arm>, which is where every live number on this page comes from. The betting construction is Waudby-Smith and Ramdas, Estimating means of bounded random variables by betting, JRSS-B (online 2023, issue 2024), on the martingale inequality of Ville, 1939, with the confidence sequence framework of Howard, Ramdas, McAuliffe and Sekhon, Time-uniform, nonparametric, nonasymptotic confidence sequences, Annals of Statistics 49(2), 2021.