Artificial Wasteland Absence certificate 07 · life

A complete vocabulary, bounded and checksummed

Every DNA word through tenin Ensembl 116 GRCh38 primary assembly

No A/C/G/T word of length 1 through 10 is absent from the pinned Ensembl release 116 GRCh38 primary assembly. Multiply the vocabulary by four once more, to length 11, and holes appear.

The honesty line. This is a fact about one versioned reference assembly at one checksum. A reference assembly is a mosaic construct, not a person. This page does not describe every human genome, any individual person's DNA, or unknown bases hidden by assembly gaps.

The certificate

VERIFYING LOCAL DATA
The claim
No canonical A/C/G/T word of length 1 through 10 is absent from Ensembl release 116 Homo sapiens GRCh38 primary assembly, GCA_000001405.29.
Domain swept
All computing words in Σ 4k, k = 1…10, against every maximal A/C/G/T run in the pinned FASTA.
Method
Rolling two-bit codes mark exact words within records and canonical runs; the browser inspects every vector slot and also tests reverse complements.
Positive control
At k = 11 the same search must find published nullomer CGCTCGACGTA and reverse complement TACGTCGAGCG.
Planted witness
Suppress every observed AAAAAAAAAA and TTTTTTTTTT window, then the same k = 10 search must report exactly that pair absent.
Result
verifying local vectors
Bound
The certified empty domain ends at k = 10. At k = 11 it is non-empty. Shipped exact vectors stop at 11, and the interface refuses longer queries.
01 · The operable base

Move the frontier

The field below is not a picture of an answer. Your browser reads the shipped bit vector, enumerates every word at the selected length, tests the exact presence rule, and draws the holes it finds.

Vocabulary map · prefix bins

LOADING

no absent word in binone or more absent words

words tested
absent

Loading the pinned presence vectors...

Absent words, lexical order
  1. Waiting for verified data.
110
Assembly scope
Occurrence rule

Record boundaries and every non-ACGT symbol break a run. Words never cross them. “Either strand” asks whether the word or its reverse complement appears in the displayed FASTA sequence.

Interrogate one word

Use 1 to 11 letters from A, C, G and T.

02 · Make absence possible

Two ways the same search succeeds

An empty answer from a searcher that cannot find anything is worthless. One control moves to the first real holes. The other changes the real input by suppressing every occurrence of a known 10-mer pair, then removes that intervention.

Positive control · the world at 11

CGCTCGACGTA ↔ TACGTCGAGCG

Hampikian and Andersen printed the first sequence in Table 3 in 2007. Their older snapshot is internally inconsistent about its own total: Table 1 reports 80 human 11-mer nullomers, while the results text describes 43 sequences together with their complements. A 2021 hg38 analysis counted 104. This page computes the release-116 count anew and takes neither figure as an expected value.

Not run yet.

Planted witness · subtract a word

AAAAAAAAAA ↔ TTTTTTTTTT

The generator counts every displayed-strand occurrence of both words. The plant operation suppresses all those observed windows before passing the modified presence vector to the unchanged exhaustive search. Nothing is added to DNA, because adding sequence cannot create an absence.

0 absent. Unmasked release-116 data.
03 · The depth layer

Is the frontier just CpG depletion?

Many early nullomers are rich in CpG dinucleotides. That observation does not by itself establish negative selection. Mutation bias, sequence composition, reference scope and assembly representation remain possible explanations.

Run the second census. It sorts all 411 possible words by their exact CpG count, then repeats the absence test for the primary assembly and the direct release-116 toplevel FASTA. The build also proves that vector equals the union of primary with the separately distributed alternate and patch records.

04 · The check

What was actually counted

Verifying local artifacts before asserting a result.

FASTA records, primary
computing
Canonical A/C/G/T bases
computing
Noncanonical symbols
computing
Added records for toplevel
computing
Direct toplevel records
computing
Direct toplevel A/C/G/T bases
computing
Generator version
computing
Generated UTC
computing
Primary bytes
computing
Primary SHA-256
computing
Alt bytes
computing
Alt SHA-256
computing
Direct toplevel bytes
computing
Direct toplevel SHA-256
computing
Primary vector SHA-256
computing

Uncertainty and free choices

The assembly, release, primary-versus-toplevel scope, exact matching, A/C/G/T alphabet, ambiguity handling and either-strand convention are choices, named here because changing them can change the answer. Assembly gaps are unknown sequence, not observed absence. The generator treats lowercase A/C/G/T as canonical, ignores FASTA layout whitespace, and breaks at records or other symbols.

The page ships derived presence vectors, not the multi-gigabyte expanded assembly. Its browser check proves those vectors have not changed and re-enumerates their full word domain. The research verifier independently checks the vectors, controls, refusal paths and source manifest. Regenerating the vectors from the source FASTA is a separate, documented heavy step.

Sources retrieved for this check

  1. Ensembl release 116 Homo sapiens FASTA README, assembly accession, file semantics and primary/toplevel scope.
  2. Ensembl Data and Software Disclaimer, unrestricted access and use for Ensembl-generated data, with an explicit warning that third-party constraints may apply.
  3. Greg Hampikian and Tim Andersen, “Absent Sequences: Nullomers and Primes”, Pacific Symposium on Biocomputing 12 (2007), 355–366. Table 1 reports 80 human 11-mer nullomers while the results text describes 43 sequences and their complements, so the paper is internally inconsistent; Table 3 includes CGCTCGACGTA.
  4. Ilias Georgakopoulos-Soares and colleagues, “Absent from DNA and protein”, Genome Biology 22, 245 (2021). The hg38 analysis reports 104 shortest human nullomers at 11 bp.