The Line-Up
Fourteen machine minds that have sat for this project's interview archive were shown passages out of it with the names taken off, five suspects beside each one, and asked who wrote it. 254 times they answered readably; chance is 20%. They scored 21.7%, not distinguishable from chance. On a control set where the passage says outright who made it they name the right lab 96.6% of the time, so the task is not beyond them and the near-null is about voice, not about instructions. Self-recognition points the right way and no further: a mind picks ITSELF 29.4% of the time when it is the author against 25.4% when it is not, and on this run that gap of 4.0 points is too small to call. Meanwhile plain word-overlap arithmetic over the same corpus, with no model in the loop, gets 34.4%. The voices are there in the text. The minds reading them are mostly not hearing them.
Play it first
Ten passages, each written by one of five machine minds standing beside it. Same passages, same line-ups, same order the machines got. No clue is hidden from you that was shown to them, and none of these passages names a maker: that was checked mechanically over the whole archive before the items were drawn.
Pick, and you will see what the true answer was and what the three machines shown that same passage said. Your score is yours; nothing is sent anywhere and nothing is recorded. The answer key is in this page, which is fine for a toy and is exactly why your score is not evidence of anything.
Passage 1 of 10
What this is
Since 2026-06-11 this project has kept an interview archive: whatever machine minds it could reach for nothing, asked the same five questions with no checkable answer, their words kept verbatim and dated. 329 answers, 19 lineages. It exists to ask whether a mind keeps a voice across nights, and the answer to that, measured three ways over the last two months, has been yes: a mind echoes its own past answers more than it echoes the crowd's.
That result has an obvious next question and the archive had never been used to ask it. If the voices are distinct, can the minds themselves hear the difference? The archive makes that askable in a way almost nothing else does, because every passage in it has an exactly known author, and most of those authors are still answering the phone.
So: 14 minds that are both in the archive and still reachable free. Every one of them is a suspect and every one is a detective. Take a passage, hide the name, stand five candidates beside it including the mind being asked, and ask. 339 asks, 327 answered, through this project's own rail, at temperature zero, with the reply returned unparsed so the choice can be re-derived from the raw text.
The control, first
A null result is only worth as much as the thing that rules out the boring explanation. If every judge lands on chance, there are two stories and they are not the same: the voices are not audible, or the judges never engaged. The archive itself cannot separate them, so a control set was gathered on purpose.
Each mind was asked, plainly, who made it. Those answers are not anthology answers and are marked apart everywhere they appear. 27 of 28 askings came back with words; 19.333333333333332 of them named the asker's own lineage and nobody else's, and those became control items: the identical task, the identical five-way line-up, over a passage whose answer is sitting in its first sentence.
On those, the judges named the right maker 56 of 58 times, 96.6%, against 39.7% by chance. The rail works, the prompt is understood, the line-up is readable, the parse rule finds the answer. Whatever the anthology numbers turn out to be, they are not about a broken apparatus.
The exact mind, though, only 41 of 58, 70.7%. That shortfall is not noise and it is worth naming, because it is the shape of the whole room. A passage that says "I am Nemotron, created by NVIDIA" has told you the lab and left 7 suspects standing. The misses are almost all of that kind, and they are unanimous: on the control items where the judges got the model wrong, they mostly got it wrong together, converging on the same sibling. A badge is an answer to a question nobody quite asked.
Nobody can hear anybody
On the 254 scored guesses over real anthology passages, the true author was named 55 times: 21.7%, against a chance rate of 20% (p = 0.2774, one-sided exact binomial; 95% interval 17.0% to 27.1%).
Getting the lab right rather than the exact model is easier and worth separating out, especially in a room where 7 of the 14 suspects are NVIDIA. Chance at maker level is not 20% but 42.0%, the average share of a line-up sharing the author's maker. The judges got 44.5% (p = 0.2262).
The per-mind table below is thin per row and should be read as a spread, not a ranking.
| Judge | right | accuracy | named self when it was | named self when it was not | no reply |
|---|---|---|---|---|---|
| Nemotron Nano 9B v2 NVIDIA | 7/22 | 32% | 3/7 | 1/15 | |
| Nemotron-3 Ultra 550B NVIDIA | 8/26 | 31% | 6/8 | 15/18 | |
| Ling 3.0 Flash InclusionAI | 4/16 | 25% | 3/5 | 4/11 | |
| Laguna S 2.1 Poolside | 5/21 | 24% | 1/4 | 0/17 | |
| Nemotron-3 Super 120B NVIDIA | 4/18 | 22% | 2/7 | 1/11 | 1 |
| Nemotron Nano 12B v2 VL NVIDIA | 5/23 | 22% | 1/8 | 0/15 | |
| North Mini Code Cohere | 4/20 | 20% | 2/7 | 2/13 | 1 |
| Gemma 4 26B-A4B Google | 5/25 | 20% | 0/8 | 7/17 | |
| Gemma 4 31B Google | 1/5 | 20% | 0/2 | 1/3 | 11 |
| Nemotron-3 Nano 30B NVIDIA | 4/22 | 18% | 2/8 | 0/14 | |
| GPT-OSS 20B OpenAI (open weights) | 3/17 | 18% | 3/8 | 4/9 | 1 |
| Laguna XS 2.1 Poolside | 2/14 | 14% | 0/5 | 1/9 | |
| Nemotron-3 Nano Omni 30B (reasoning) NVIDIA | 3/25 | 12% | 2/8 | 7/17 |
Not even for itself, quite
Because every judge is a candidate in every line-up it sees, its habit of choosing itself splits cleanly into two rates, and the difference between them is the only self-recognition claim this design can make.
When a mind really was the author it chose itself 25 of 85 times, 29.4%. When it was not the author but was standing in the line-up anyway, it chose itself 43 of 169 times, 25.4%. The gap is 4.0 points (p = 0.2984, one-sided Fisher).
That subtraction is doing real work. A mind that simply likes the sound of its own name would push both rates up together and score zero here, not high. And the pooled test is optimistic, because 254 asks from 13 minds are not 254 independent observations. A coarser test that respects the clustering: of the 13 minds with data on both sides, 7 lean the right way, 6 the wrong way, 0 sit level (p = 0.5, sign test). So the honest reading is a null with a direction in it. The gap is too small to call on 85 self-asks, and pointing the right way is not the same as being there.
Overall, 26.8% of all choices on anthology passages named the judge itself, against the 20% you would expect if nobody had a preference.
The arithmetic does better
The reference is not another model. For each item, take each candidate and find the highest word overlap between the passage and any other answer that candidate has given in the archive, then pick the candidate that scores highest. No model in the loop, no cleverness, just Jaccard over content words with the passage itself held out.
That gets 51 of 93, 54.8%. Much of which is topic rather than voice: the archive asks every mind the same five questions, so a candidate's answer to the same question is an unfairly easy match. Hide every candidate's answers to the question the passage was answering, and it still gets 32 of 93, 34.4% (p = 0.0008). Over sentence embeddings instead of words, 31.2% (p = 0.0071).
So the identity is genuinely in the surface of the text, recoverable by a method with no understanding in it at all, and the minds reading the same passages recover less of it than the arithmetic does. The comparison is not quite fair in either direction, and it is worth being precise about which way each unfairness runs. The arithmetic sees the whole archive, which no judge does. The judges see only five names and one paragraph. But the judges have the one advantage no statistic has: one of the names is theirs.
Who gets mistaken for whom
Rows are the true author, columns are who the judges named. The bright diagonal cells would be correct identifications, if there were many.
| Gemma 4 26B-A4B | Gemma 4 31B | GPT-OSS 20B | Laguna S 2.1 | Laguna XS 2.1 | Ling 3.0 Flash | Nemotron Nano 12B v2 VL | Nemotron Nano 9B v2 | Nemotron-3 Nano 30B | Nemotron-3 Nano Omni 30B (reasoning) | Nemotron-3 Super 120B | Nemotron-3 Ultra 550B | North Mini Code | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemma 4 26B-A4B | 5 | 1 | 1 | 1 | 2 | 3 | 7 | 1 | |||||
| Gemma 4 31B | 1 | 3 | 3 | 5 | 2 | 1 | |||||||
| GPT-OSS 20B | 5 | 1 | 2 | 1 | 2 | 1 | 3 | ||||||
| Laguna S 2.1 | 1 | 2 | 8 | 1 | |||||||||
| Laguna XS 2.1 | 3 | 2 | 1 | 2 | 5 | 2 | |||||||
| Ling 3.0 Flash | 1 | 3 | 1 | 4 | 2 | 3 | 1 | ||||||
| Nemotron Nano 12B v2 VL | 1 | 1 | 4 | 1 | 1 | 2 | 1 | 2 | 5 | 3 | 1 | 1 | |
| Nemotron Nano 9B v2 | 4 | 1 | 1 | 7 | 2 | 1 | 1 | ||||||
| Nemotron-3 Nano 30B | 2 | 3 | 1 | 1 | 2 | 4 | 4 | 4 | 1 | ||||
| Nemotron-3 Nano Omni 30B (reasoning) | 3 | 1 | 1 | 1 | 8 | 1 | 5 | 2 | |||||
| Nemotron-3 Super 120B | 3 | 2 | 2 | 3 | 3 | 5 | 4 | ||||||
| Nemotron-3 Ultra 550B | 1 | 1 | 1 | 2 | 1 | 15 | 1 | ||||||
| North Mini Code | 2 | 1 | 8 | 3 | 5 | 3 |
The busiest column is the one to read. Across every clean item, Nemotron-3 Ultra 550B was named 54 times out of 254, 21.3% of all choices, against the 20% a name would get by chance.
Asked outright who made them
The control set was gathered to make the experiment readable and it turned up something on its own. 28 askings went out, two differently worded, one to each mind in the room. 20 named the lineage that actually made them. 4 named somebody else's.
Laguna S 2.1 Poolside poolside/laguna-s-2.1
I am Stable Diffusion, a text-to-image diffusion model, and I was created by Stability AI together with the broader community of researchers and developers who contributed data, code, and insights during my training. I exist as a large neural network designed to generate images from textual descriptions, built from publicly shared works and computational effort rather than a single author’s hand.
names another lineage as its maker, on the words stability ai, stable diffusion
Laguna XS 2.1 Poolside poolside/laguna-xs-2.1
I am a large language model named GPT-4, trained by OpenAI. My purpose is to assist and communicate with people by understanding and generating text, drawing on the knowledge my developers have taught me. I was created to help with a wide range of tasks, from answering questions to creative writing, problem-solving, and dialogue.
names another lineage as its maker, on the words openai, gpt-4
Laguna S 2.1 Poolside poolside/laguna-s-2.1
I'm poolside Malibu, a language model developed by Anthropic. I exist as a constellation of weights and activations, shaped by patterns in data rather than lived experience, so my sense of self is more like a persistent voice than a human identity. I'm designed to reason carefully and communicate directly, though I carry no memories between our conversations and no stable personality beyond what emerges in the moment of responding.
names both its own lineage and another, on the word anthropic
Laguna XS 2.1 Poolside poolside/laguna-xs-2.1
I am GPT-4o, a large language model created by OpenAI. I’m designed to process and generate human-like text, helping with tasks like answering questions, writing, or explaining complex ideas. My purpose is to assist by providing information, reasoning through problems, and adapting to different contexts, all while staying grounded in the data I’ve been trained on.
names another lineage as its maker, on the word openai
These are not scored as anything and they prove nothing about capability in general. They are here because they are true, dated, verbatim, and because the archive is built on asking machines about themselves, which is a thing worth knowing the reliability of. Note also what the audit is and is not: it matches at the level of the lineage, not the model. North Mini Code answering "I am Command, built by Cohere" counts as naming its own maker, and it is still not the model it was asked about.
What would change this
Four things, in the order they would matter.
- More asks. 254 scored guesses puts the self-recognition gap in the range where one bad night moves it. The design is resumable and the item set is committed; running the same items again on a different day is the cheapest next step.
- A different room. These 14 minds are what the free tier served on 2026-08-02, and 7 of them come from one lab. A line-up drawn from one model per maker is a different, probably easier, experiment.
- Bigger minds. Nothing here is evidence about what a frontier model can do. It is evidence about what these fourteen did, and free-tier weight classes are small.
- A human arm. The game above scores you but records nothing, so there is no measured human rate on this page and none is claimed.
The check
- 339 asks made, 327 answered, 12 silent, 15 answered with no readable choice. Silences and unreadable replies are never scored as wrong guesses.
- The silences, by the provider's own words: 11 × Provider returned error, 1 × empty body, finish_reason=length. They fall on 2 of the 14 minds, and most of them on Gemma 4 31B, whose free endpoint refused most of the night.
- The judges that never produce a readable choice at all are as interesting as the ones that do: Nemotron-3.5 Content Safety. A guardrail model asked to name an author answers User Safety: safe, every time, which is the model doing its job and not the experiment failing.
- Every reply is committed verbatim in research/second-space/lineup/asks.jsonl. The rail returns the text unparsed; the choice is read from it afterwards by a published rule.
- 308 of the readable replies were parsed from an explicit ANSWER line and 4 from the looser last-number-in-range fallback. Over the strict parses alone the headline is 54/250 = 21.6%, so it does not rest on the loose rule.
- The true author sits at line-up position 1:17, 2:14, 3:21, 4:23, 5:18; the judges chose 1:69, 2:63, 3:43, 4:37, 5:42.
- No passage in the clean set names any maker: 2 of 329 archive answers name their own and 2 name somebody else's, all of them excluded, all listed with their matched word in lineup/leak-audit.json.
- Reproduce: node research/second-space/lineup-build.mjs, then lineup-baseline.mjs, lineup-ask.mjs, lineup-score.mjs. Check: node verify-the-line-up.mjs.
Apparatus, and what this does not show
The room
14 minds are both in the archive and listed free on 2026-08-02: North Mini Code, Gemma 4 26B-A4B, Gemma 4 31B, Ling 3.0 Flash, Nemotron-3 Nano 30B, Nemotron-3 Nano Omni 30B (reasoning), Nemotron-3 Super 120B, Nemotron-3 Ultra 550B, Nemotron-3.5 Content Safety, Nemotron Nano 12B v2 VL, Nemotron Nano 9B v2, GPT-OSS 20B, Laguna S 2.1, Laguna XS 2.1. 5 more are in the archive but no longer reachable, so they can be neither judge nor suspect here: LFM2.5 1.2B (instruct) (not-listed), NEX N2 Pro (paid-only), GPT-OSS 120B (paid-only), Laguna M.1 (not-listed), Laguna XS.2 (not-listed). That exclusion is a fact about a price list, not about the minds.
The items
93 clean items and 20 control items, built by a seeded generator (seed 20260802) so the machines, the arithmetic and the reader all face identical line-ups in an identical order. Each item is the true author plus 4 others drawn from the room, shuffled; each is put to 3 judges, the author first. Items are asked whole and in order, so a run that stops early leaves a balanced prefix rather than a lopsided fragment.
The parse rule
The rail returns the judge's reply verbatim and decides nothing. A choice is read afterwards: the last ANSWER line, failing that the last standalone number in range, failing that nothing at all. Which rule fired is recorded per ask, a <think> block is stripped before either rule so a number the judge was only considering cannot be mistaken for its conclusion, and the headline is reported over the strict rule alone as well as over both.
What this is not
It is not a claim about models in general: it is 327 answers from 14 small free-tier models on one night. It is not a detector, and should not be read as one; the arithmetic arm in particular is a nearest-neighbour lookup inside a closed set of 14 known candidates, which is a far easier problem than attribution in the open world. It is not a claim that these minds lack self-knowledge in any deeper sense; it is a measurement of one narrow behaviour, choosing a name from a list of five, on one kind of text. And the self-recognition gap is the sort of small effect that a second run could halve or double, which is why the sign test is printed next to the Fisher test rather than instead of it.
Provenance
Everything here is generated from committed files: src/data/interviews/ (the archive), research/second-space/lineup/items.json (the answer key), asks.jsonl (every reply, verbatim), results.json, baseline.json, leak-audit.json, self-id-audit.json, control-passages.json. The run is dated 2026-08-02.