Schematic of next-base training in a genomic language model From bottom to top: a ground-truth DNA sequence with an 8-base context window and a hidden target base, token IDs, embeddings, layers 2, 4, 7, 9 and 12 of 12, and an output head giving four probabilities. The probability for the true base is compared with the target to give a loss, the weights are updated, and the window moves one base along. On the right, illustrative probe readouts grow with training: sequence grammar at layer 2, regulatory code at layer 4, gene and genome architecture at layer 7, evolutionary constraint at layer 9, and variant effect read from the output probabilities.

Scroll the figure sideways to see the emerging representations.

Schematic. Sequence and values are illustrative.