Introducing LOAM: Learning Biology from Long-Read Soil Metagenomes

A family of genomic language models trained on environmental DNA

Tim Rajakumar, Quentin Ferry
1 October 2026
6 min

Today we introduce LOAM, a family of genomic language models trained exclusively on long-read environmental metagenomes.

Across a controlled series ranging from 25 million to 624 million parameters, LOAM learns biologically informative representations, uses genomic context across several kilobases, and performs competitively with much larger models on selected prokaryotic benchmarks.

Learning directly from soil genomes

Soil and other environmental microbiomes contain enormous genomic diversity. Their microorganisms have spent billions of years adapting to complex ecological environments, evolving genes, proteins, enzymes, biosynthetic pathways and other biological functions that remain only partially characterised.

Metagenomics increasingly gives us access to this diversity. Conventional short-read sequencing generates large numbers of small DNA fragments that must be computationally reassembled, often leaving complex microbial genomes fragmented. Long-read sequencing spans much larger stretches of DNA, producing more contiguous assemblies and preserving longer genomic neighbourhoods around genes and other functional elements.

But accessing this sequence diversity creates another challenge. Environmental genomes contain millions of genes and other sequence elements whose functions are only partially characterised. Experimental annotation cannot keep pace with the amount of sequence becoming available.

Could models that learn directly from DNA, without the need for human annotations, help extract useful biological information from this largely unexplored sequence space?

Training LOAM

Genomic language modelling is already a rapidly developing field. Models including the Evo family, Nucleotide Transformer and gLM2 have shown that self-supervised learning across large collections of DNA sequences can produce representations that capture biologically useful information. Existing genomic language models have been trained across a range of reference-genome and metagenomic datasets. We were interested in a more specific question: Could a genomic language model trained directly on long-read, environmental metagenomes learn biologically useful representations from this microbial diversity? To test this, we developed LOAM.

We trained LOAM on the public Micro Flora Danica long-read dataset, using 15,640 species-level representative microbial genomes reconstructed from Oxford Nanopore sequencing of soil, sediment and water environments across Denmark. The resulting pre-training corpus contained approximately 67.5 billion DNA bases. An important feature of this corpus is its genomic continuity. Long-read assemblies provide longer uninterrupted stretches of sequence during training, exposing the models to more complete genomic neighbourhoods rather than repeatedly reaching the boundaries of fragmented contigs.

How next-base training works in a genomic language model. Generic schematic, not LOAM's architecture; sequence, layer assignments and values are illustrative, not measured probe results.

LOAM uses the same broad self-supervised learning principle that underpins autoregressive language models, but its input is DNA rather than text. The model reads nucleotide sequence and repeatedly predicts what comes next. The objective is simple, but performing it accurately across billions of bases requires the model to learn recurring sequence patterns and higher-order structure within genomic DNA. This is particularly attractive for environmental microbiology because much of its sequence diversity remains poorly characterised. A model can begin learning from genomic sequence even where experimentally validated functional annotations are sparse.

The important question is therefore not simply whether LOAM becomes good at predicting DNA. It is whether learning to predict environmental DNA produces representations that capture useful biological information.

Does LOAM learn useful biology?

We evaluated LOAM across three biological questions designed to probe different aspects of its representations.

For gene essentiality and enzyme function, we froze the pretrained models and trained simple linear classifiers on their internal representations. For mutation effects, we used the models' own sequence likelihoods directly, without fitting predictors to the experimental outcomes.

Benchmark scores for thirteen genomic language models, 16 September run. Embedding tasks use each model's best layer; mutation effects are scored zero-shot.

Which genes are essential?

Using the BacBench gene-essentiality benchmark, we asked whether LOAM representations could distinguish experimentally labelled essential genes from non-essential genes. Across the parameter range tested, LOAM outperformed the comparably sized genomic language models included in our comparison. The much larger 7-billion-parameter Evo models remained strongest overall.

What does an enzyme do?

We next tested whether LOAM representations contained information about enzyme function. Using a 128-class DNA-based enzyme-function benchmark from DGEB, the largest LOAM model achieved the highest performance among the models tested. Importantly, LOAM had not been trained to recognise enzyme classes. Information relevant to function emerged from learning genomic sequence itself.

How do sequence variants affect biological function?

Finally, we evaluated LOAM on 11 bacterial deep-mutational-scanning experiments from RNAGym. These experiments measure the effects of sequence variants across phenotypes including antibiotic resistance, cellular fitness, molecular transport, enzyme activity and thermostability. Each mutation was scored using only the change in sequence likelihood predicted by LOAM. No predictor was trained on the experimental labels.

Performance improved consistently with model size. Our flagship LOAM-624M achieved aggregate performance comparable to the much larger Evo 1.5 model. Evo2-7B remained strongest overall, while LOAM-624M outperformed it on four of the eleven individual assays.

Four things we learned from LOAM

The benchmarks tell us that useful biological signal exists in LOAM's representations. Because all four LOAM models were trained on the same data in the same order, the series also provides a controlled system for asking how genomic representations change with scale, context and training-corpus composition.

1.Larger models make better use of more sequence

Across a single pass through the same training corpus, larger models predicted held-out genomic sequence more accurately. Additional training sequence improved every model, with larger models benefiting more strongly than smaller ones. None of the training curves showed clear saturation, suggesting that additional training sequence could yield further gains.

2.Genomic context changes predictions thousands of bases downstream

Long context is only useful if a model actually uses it. We therefore designed an experiment to test this. We held the sequence being predicted constant while changing the long-range genomic context that preceded it. The model received either its true upstream neighbourhood, no additional upstream sequence, or unrelated sequence from another organism.

Correct context helped. Foreign context hurt.

How much DNA fits in one LOAM window?. Log-scale chart of DNA lengths against LOAM's context window of 8,192 bp. Nucleotide: 1 bp (dot at 1 bp); Codon: 3 bp (dot at 3 bp); Regulatory motif: 6–20 bp (dot at 12 bp, range 6 to 20 bp, literature range); Intergenic DNA: 10s–100s bp (dot at 100 bp, range 4 to 461 bp, 5th-95th percentile); Protein-coding gene: ~1 kb (dot at 1,000 bp, range 186 to 2,139 bp, 5th-95th percentile); Operon: a few kb (dot at 3,000 bp, range 1,000 to 8,192 bp, 5th-95th percentile); Genomic neighborhood: 10–20 kb (dot at 15,000 bp, range 10,000 to 20,000 bp, literature range); Biosynthetic gene cluster: ~40 kb (dot at 41,000 bp, range 3,800 to 101,645 bp, 5th-95th percentile); Whole genome: a few Mb (dot at 4,000,000 bp, range 1,636,530 to 7,535,756 bp, 5th-95th percentile). Dots mark illustrative sizes; grey bars show published ranges.Effect of distant upstream DNA, by model size. Change in loss at positions 3,072–4,095 of a block when its 4 kb of upstream DNA is real or foreign, compared with none. Log scales. Foreign DNA raises the loss: LOAM-25M 0.0055 (95% interval 0.0054 to 0.0057); LOAM-100M 0.0076 (95% interval 0.0074 to 0.0078); LOAM-340M 0.0064 (95% interval 0.0063 to 0.0066); LOAM-624M 0.0064 (95% interval 0.0063 to 0.0066). Real DNA lowers the loss: LOAM-25M 0.0006 (95% interval 0.0006 to 0.0007); LOAM-100M 0.0009 (95% interval 0.0008 to 0.0009); LOAM-340M 0.0009 (95% interval 0.0008 to 0.0009); LOAM-624M 0.0010 (95% interval 0.0009 to 0.0011). Bars are 95% bootstrap intervals (12,000 windows); most are smaller than the dots.

The advantage of the true genomic neighbourhood was strongest nearby but remained detectable several thousand nucleotides downstream. Larger models generally extracted greater predictive value from this distal context. LOAM is therefore not simply reacting to the nucleotides immediately adjacent to the position being predicted. Information from the wider genomic neighbourhood changes its expectations.

3.The final layer is not always the most biologically informative

We also asked where within the models biological information was most accessible. Instead of probing only their final representations, we trained linear probes from every hidden layer. Across LOAM and the comparator models, intermediate layers outperformed the final layer in 19 of 26 model-task comparisons. The representation best suited to next-nucleotide prediction is therefore not necessarily the representation from which another biological property can most easily be read.

4.Biological coverage matters, not just dataset size

The mutation experiments revealed another important pattern. LOAM's performance relative to Evo2 differed considerably between target genes. We therefore asked whether these differences related to how strongly the relevant biology was represented in LOAM's training corpus.

Prediction gap and training homologues. Each dot is an RNAGym assay; right means more detectable homologues. Vertical axis: Prediction gap (LOAM − Evo2, Spearman ρ). Horizontal axis: Detected homologues (searched contigs), log scale. BCHB: 0 homologues, gap -0.205 (95% interval -0.270 to -0.139); PSAE: 50 homologues, gap -0.203 (95% interval -0.268 to -0.137); CCDB: 65 homologues, gap -0.055 (95% interval -0.135 to +0.025); F7YBW8: 399 homologues, gap +0.132 (95% interval +0.107 to +0.158); ESTA: 917 homologues, gap -0.101 (95% interval -0.157 to -0.045); BLAT ’14: 1498 homologues, gap -0.006 (95% interval -0.040 to +0.028); BLAT ’13: 1498 homologues, gap -0.035 (95% interval -0.110 to +0.040); MLAC: 1773 homologues, gap -0.071 (95% interval -0.112 to -0.030); RNC: 12561 homologues, gap +0.076 (95% interval +0.041 to +0.112); IF1: 14883 homologues, gap +0.063 (95% interval -0.006 to +0.132); Q837P4 ≥: 100000 homologues, gap +0.134 (95% interval +0.045 to +0.223). LOAM ahead on 4 assays, Evo2 ahead on 7. Dot area reflects gene length; bars are conservative 95% gap intervals.

LOAM tended to perform relatively better when the wild-type sequence was less surprising to the model and when more detectable homologues of the target gene occurred in the training data. Its largest deficits tended to occur where related biology was scarce or absent. The relationship was not absolute, and corpus representation did not explain every performance difference.

But it points towards an important principle: What matters is not only how much DNA a genomic language model sees, but which biology that DNA represents.

That observation is particularly important for environmental genomics. It suggests that future genomic models may benefit not simply from ever larger sequence collections, but from deliberately expanding their exposure to biologically relevant and underrepresented parts of sequence space.

For Soilytix, it also raises a broader question: if the biology represented in a model matters, can ecological observations help us decide where in nature to search for the biology we want? That question takes us beyond the model itself and towards the broader discovery system we are building at Soilytix. In our companion post, we explore how field observations could be combined with environmental genomics and computational models to focus the search for new biological crop-protection actives.

Built to be built on

Scientific progress compounds when others can inspect, test and build upon it. That is why we are releasing all four LOAM models on Hugging Face for research use alongside the accompanying manuscript. Releasing the complete series provides more than a single checkpoint. It gives researchers a controlled model family for studying how genomic representations, context utilisation and biological information change with scale.

About the authors

Tim Rajakumar

Co-founder, Chief Scientific Officer

Tim co-founded Soilytix and sets its research direction, owning the science behind its assays and claims. He leads the company's applied science and co-leads its AI work on bioassets with Quentin.

Quentin Ferry

Co-founder, Chief AI Scientist

Quentin co-founded Soilytix and leads its AI research, building deep learning and foundation models to better mine and understand soil biology. That work includes LOAM, the family of genomic language models introduced in this post.

Follow LOAM

One or two short notes a month. Unsubscribe any time. See our privacy notice.