Life, counted by its probabilities
The Receptors We Meet Twice
Explore a declared V(D)J toy model through event-key probability and entropy, then draw two independent toy repertoires and see hundreds of shared synthetic event keys where a uniform 1014 sequence space predicts almost none.
A repertoire is not an urn filled with every possible receptor once. It is a weighted machine. Some sequences sit near the mouth; most occupy a tail so thin that no person will ever draw them.
Assemble one receptor
One click samples V, D, J, four deletion choices, two insertion lengths, and every inserted base.
The nucleotide string is assembled from synthetic terminal strings. The 48 V, 2 D, and 13 J counts are an approximate IMGT functional-gene scale, not an allele catalogue. Pkey is the probability of the sampled event-description key after summing 32 explicit latent aliases. It is not Pgen of the displayed nucleotide string, because other event descriptions that render the same string are not collapsed.
The first multiplication is small: 48 × 2 × 13 = computing. The enormous event-key support comes later, when junctions can delete bases and insert variable strings. Because different descriptions can render the same nucleotide string, this support is an upper bound on distinct emitted strings.
These are a spread of literature values for different chains, stages, and definitions, not an uncertainty interval or settings of this toy. Murugan et al. report about 52 bits for TCR-beta rearrangement events and about 47 bits for nucleotide CDR3 sequences after convergent events are summed. Mora and Walczak's review separately characterizes TCR-beta generation as about 43 bits. The modelled post-selection estimate is about 38 bits, while TCR-alpha and immunoglobulin heavy chain occupy different scales near 30 and 70 bits.
The probability landscape
Twenty thousand draws from the declared model, binned by −log10 Pkey. Height is sample count, not biological abundance.
Activate the chart with Enter, Space, click, or tap to rerun it. The empirical literature reports roughly 20 orders of sequence Pgen variation. This toy instead measures event-key Pkey and labels its sampled span accordingly.
Event-key support asks, “Can the model generate this full description?” Event-key entropy asks, “How surprised should I be by the description drawn?” Neither is the exact inventory or entropy of emitted nucleotide strings.
The collision that should not happen
Now take the sophisticated objection seriously. Even an effective sequence space near 1014 is larger than one person's sampled repertoire. So perhaps the diversity number is only academic. The objection opens the useful door.
If every sequence were equally likely, two samples of size N would share about N²/S sequences when N is small beside the space S. At N = 106 and S = 1014, that is 0.01. The toy comparison below counts matching event-description hashes instead. Since convergent event descriptions are not collapsed, that count is a lower bound on sharing of the rendered nucleotide strings.
Draw two independent toy repertoires
The two seeded streams share no simulated clonal ancestry, antigen history, or random state. They share only the same event distribution.
Event-hash lower bound / uniform sequence null: ready
This is a simulation of the shipped event-key distribution, not a claim about the exact sharing of a real pair. Different event descriptions that render the same nucleotide string are not collapsed, so event-hash sharing is a lower bound on rendered-sequence sharing. It is not a count of nucleotide clonotypes or a biological publicness estimate.
Elhanati and colleagues found that a data-driven recombination model, plus a simple selection factor, predicts the observed sharing spectrum between unrelated people. A uniform model misses by orders of magnitude. Convergent recombination is the leading explanation for publicness in that analysis. No shared clonal ancestry or shared antigen exposure is needed for the basic excess expected from independent generation.
The check
Every green model number below is recomputed in this page from the distributions in its script. The standalone verifier repeats the model independently with the same fixed seeds.
Observed or published
TCR-beta generation is described near 43 bits in review prose and at 47 bits for nucleotide CDR3 sequences in Murugan et al.; rearrangement events are 52 bits there. Post-selection is about 38 bits, TCR-alpha about 30, and IGH about 70.
Diagnostic inference
Sequence entropy is below sequence support when generation probabilities are non-uniform. Public clonotypes are sequence sharing above a uniform null, explained quantitatively by generation bias plus selection in the cited work.
Toy event-key output
Sample space: hashes of full event descriptions, including gene choices, deletions, insertion lengths, and inserted bases. computing segment combinations. computing capped event-key support, an upper bound on nucleotide-string support. computing event-key entropy after marginalizing the artificial aliases.
Free choices: 48/2/13 approximate segment counts; exponential gene weights; four deletion distributions; base frequencies; a geometric-like insertion stop probability of 0.17; a 22-base cap per junction; 32 equiprobable latent aliases; fixed seeds; and the uniform sequence comparison space of 1014. None is fitted here. Convergent event descriptions beyond the artificial aliases are not collapsed.
What is measured, what is modelled, and what is narrowed
Measured in the cited work. Human repertoire sequence data support non-uniform recombination models, receptor-specific Pgen, convergent recombination, entropy estimates that differ by receptor type and selection stage, and sharing spectra between unrelated people.
Computed exactly here. The component event-key entropies, event-key support, Pkey for every sampled description, seeded histogram, direct shared event-hash count, and uniform sequence null all follow from the displayed model. The capped event-key support is finite only to make its upper bound on nucleotide-string support explicit. Allowing longer insertions would enlarge it without forcing entropy to grow in step.
Narrowed. The page does not compute sequence Pgen by summing every event description that renders a nucleotide string, and it does not reproduce Murugan's maximum-likelihood inference or Elhanati's cohort analysis. Those require sequence-level aggregation or the original data and pipelines. The toy demonstrates non-uniform event generation with declared parameters. Synthetic terminal strings are used instead of presenting invented sequences as IMGT allele records.
Segment bookkeeping. V, D, and J counts vary with haplotype, ancestry, and whether functional genes, open reading frames, and pseudogenes are included. The chosen counts are a scale model. The product 1,248 is not the source of a 1030 support. Junctional strings are.
Historical lineages. Perelson and Oster's 1979 paper concerns how a finite repertoire can cover antigen shape space reliably. Modern generation entropy and public-clonotype calculations come from Murugan, Mora, Walczak, Callan, Elhanati, and colleagues.