Today we introduce LOAM, a family of genomic language models trained exclusively on long-read environmental metagenomes.
Across a controlled series ranging from 25 million to 624 million parameters, LOAM learns biologically informative representations, uses genomic context across several kilobases, and performs competitively with much larger models on selected prokaryotic benchmarks.
Learning directly from soil genomes
Soil and other environmental microbiomes contain enormous genomic diversity. Their microorganisms have spent billions of years adapting to complex ecological environments, evolving genes, proteins, enzymes, biosynthetic pathways and other biological functions that remain only partially characterised.
Metagenomics increasingly gives us access to this diversity. Conventional short-read sequencing generates large numbers of small DNA fragments that must be computationally reassembled, often leaving complex microbial genomes fragmented. Long-read sequencing spans much larger stretches of DNA, producing more contiguous assemblies and preserving longer genomic neighbourhoods around genes and other functional elements.
But accessing this sequence diversity creates another challenge. Environmental genomes contain millions of genes and other sequence elements whose functions are only partially characterised. Experimental annotation cannot keep pace with the amount of sequence becoming available.
Could models that learn directly from DNA, without the need for human annotations, help extract useful biological information from this largely unexplored sequence space?
Training LOAM
Genomic language modelling is already a rapidly developing field. Models including the Evo family, Nucleotide Transformer and gLM2 have shown that self-supervised learning across large collections of DNA sequences can produce representations that capture biologically useful information. Existing genomic language models have been trained across a range of reference-genome and metagenomic datasets. We were interested in a more specific question: Could a genomic language model trained directly on long-read, environmental metagenomes learn biologically useful representations from this microbial diversity? To test this, we developed LOAM.
We trained LOAM on the public Micro Flora Danica long-read dataset, using 15,640 species-level representative microbial genomes reconstructed from Oxford Nanopore sequencing of soil, sediment and water environments across Denmark. The resulting pre-training corpus contained approximately 67.5 billion DNA bases. An important feature of this corpus is its genomic continuity. Long-read assemblies provide longer uninterrupted stretches of sequence during training, exposing the models to more complete genomic neighbourhoods rather than repeatedly reaching the boundaries of fragmented contigs.
LOAM uses the same broad self-supervised learning principle that underpins autoregressive language models, but its input is DNA rather than text. The model reads nucleotide sequence and repeatedly predicts what comes next. The objective is simple, but performing it accurately across billions of bases requires the model to learn recurring sequence patterns and higher-order structure within genomic DNA. This is particularly attractive for environmental microbiology because much of its sequence diversity remains poorly characterised. A model can begin learning from genomic sequence even where experimentally validated functional annotations are sparse.
The important question is therefore not simply whether LOAM becomes good at predicting DNA. It is whether learning to predict environmental DNA produces representations that capture useful biological information.
Does LOAM learn useful biology?
We evaluated LOAM across three biological questions designed to probe different aspects of its representations.
For gene essentiality and enzyme function, we froze the pretrained models and trained simple linear classifiers on their internal representations. For mutation effects, we used the models' own sequence likelihoods directly, without fitting predictors to the experimental outcomes.
Which genes are essential?
Using the BacBench gene-essentiality benchmark, we asked whether LOAM representations could distinguish experimentally labelled essential genes from non-essential genes. Across the parameter range tested, LOAM outperformed the comparably sized genomic language models included in our comparison. The much larger 7-billion-parameter Evo models remained strongest overall.
What does an enzyme do?
We next tested whether LOAM representations contained information about enzyme function. Using a 128-class DNA-based enzyme-function benchmark from DGEB, the largest LOAM model achieved the highest performance among the models tested. Importantly, LOAM had not been trained to recognise enzyme classes. Information relevant to function emerged from learning genomic sequence itself.
How do sequence variants affect biological function?
Finally, we evaluated LOAM on 11 bacterial deep-mutational-scanning experiments from RNAGym. These experiments measure the effects of sequence variants across phenotypes including antibiotic resistance, cellular fitness, molecular transport, enzyme activity and thermostability. Each mutation was scored using only the change in sequence likelihood predicted by LOAM. No predictor was trained on the experimental labels.
Performance improved consistently with model size. Our flagship LOAM-624M achieved aggregate performance comparable to the much larger Evo 1.5 model. Evo2-7B remained strongest overall, while LOAM-624M outperformed it on four of the eleven individual assays.
Four things we learned from LOAM
The benchmarks tell us that useful biological signal exists in LOAM's representations. Because all four LOAM models were trained on the same data in the same order, the series also provides a controlled system for asking how genomic representations change with scale, context and training-corpus composition.
1.Larger models make better use of more sequence
Across a single pass through the same training corpus, larger models predicted held-out genomic sequence more accurately. Additional training sequence improved every model, with larger models benefiting more strongly than smaller ones. None of the training curves showed clear saturation, suggesting that additional training sequence could yield further gains.
2.Genomic context changes predictions thousands of bases downstream
Long context is only useful if a model actually uses it. We therefore designed an experiment to test this. We held the sequence being predicted constant while changing the long-range genomic context that preceded it. The model received either its true upstream neighbourhood, no additional upstream sequence, or unrelated sequence from another organism.
Correct context helped. Foreign context hurt.
The advantage of the true genomic neighbourhood was strongest nearby but remained detectable several thousand nucleotides downstream. Larger models generally extracted greater predictive value from this distal context. LOAM is therefore not simply reacting to the nucleotides immediately adjacent to the position being predicted. Information from the wider genomic neighbourhood changes its expectations.
3.The final layer is not always the most biologically informative
We also asked where within the models biological information was most accessible. Instead of probing only their final representations, we trained linear probes from every hidden layer. Across LOAM and the comparator models, intermediate layers outperformed the final layer in 19 of 26 model-task comparisons. The representation best suited to next-nucleotide prediction is therefore not necessarily the representation from which another biological property can most easily be read.
4.Biological coverage matters, not just dataset size
The mutation experiments revealed another important pattern. LOAM's performance relative to Evo2 differed considerably between target genes. We therefore asked whether these differences related to how strongly the relevant biology was represented in LOAM's training corpus.
LOAM tended to perform relatively better when the wild-type sequence was less surprising to the model and when more detectable homologues of the target gene occurred in the training data. Its largest deficits tended to occur where related biology was scarce or absent. The relationship was not absolute, and corpus representation did not explain every performance difference.
But it points towards an important principle: What matters is not only how much DNA a genomic language model sees, but which biology that DNA represents.
That observation is particularly important for environmental genomics. It suggests that future genomic models may benefit not simply from ever larger sequence collections, but from deliberately expanding their exposure to biologically relevant and underrepresented parts of sequence space.
For Soilytix, it also raises a broader question: if the biology represented in a model matters, can ecological observations help us decide where in nature to search for the biology we want? That question takes us beyond the model itself and towards the broader discovery system we are building at Soilytix. In our companion post, we explore how field observations could be combined with environmental genomics and computational models to focus the search for new biological crop-protection actives.
Built to be built on
Scientific progress compounds when others can inspect, test and build upon it. That is why we are releasing all four LOAM models on Hugging Face for research use alongside the accompanying manuscript. Releasing the complete series provides more than a single checkpoint. It gives researchers a controlled model family for studying how genomic representations, context utilisation and biological information change with scale.
About the authors

Tim Rajakumar
Co-founder, Chief Scientific OfficerTim co-founded Soilytix and sets its research direction, owning the science behind its assays and claims. He leads the company's applied science and co-leads its AI work on bioassets with Quentin.

Quentin Ferry
Co-founder, Chief AI ScientistQuentin co-founded Soilytix and leads its AI research, building deep learning and foundation models to better mine and understand soil biology. That work includes LOAM, the family of genomic language models introduced in this post.
