Crop protection is under pressure. Pathogens are becoming resistant, and regulators are taking established products off the market. The industry needs new solutions, and that means searching new sources.
Most discovery still draws on the same well-known microbes and chemistry. Soil is one of the largest sources that remains largely unexplored.
It is one of the most biologically diverse places on Earth. Its microbes have spent billions of years competing with each other and living alongside plants. Their DNA holds the recipes for the proteins, enzymes and natural compounds they use to fight, feed and survive.
Most of it has never been explored. That makes soil one of the richest untapped sources for the next generation of biological crop protection.
The real challenge: knowing where to look
Long-read sequencing now lets us read the DNA of soil microbes straight from a sample, without first growing them in a lab. Getting the DNA is no longer the hard part.
Choosing what to do with it is. A handful of soil holds thousands of species and millions of genes. Every candidate has to be tested in a lab, and lab testing is slow and expensive. You cannot test everything. The real challenge is choosing the few worth testing.
Our approach: let the field decide. Nature has already run the experiment, in real fields, under real conditions. We read the result.
Nature already ran the experiment
Picture two fields where the same crop pathogen is present.

In one, the crop gets sick. In the other, the pathogen is there, but the disease never appears.
The second field may contain microbes that help keep the disease in check.
Microbes are not the only possible explanation. Crop variety, weather, management, soil chemistry, how much of the pathogen is present and other factors can all influence whether disease develops.
But across enough fields, differences in disease outcome can point to something useful. They let us ask a simple question: which microbes keep turning up where the disease stays away?
Soilytix has built a dataset spanning thousands of agricultural field samples, combining microbial community profiles with information including pathogen occurrence and observed disease outcomes.
That changes where discovery starts. Instead of screening microbes picked with little knowledge of where they came from, we can start with the ones linked to the result that matters: a healthy crop despite the pathogen.
The next question: which of their genes are worth testing?
Reading genes nobody has studied
Sequencing gives us the DNA. It does not tell us what that DNA does. That is where LOAM comes in.
Most genes in soil have never been studied, so nobody knows what they do. That is exactly where new solutions can hide, and exactly where traditional tools struggle, because they rely on comparing a gene with genes that are already known. LOAM, our family of AI models, helps us judge which of those unknown genes are worth a closer look.
LOAM learns the way people learn a language: not from a dictionary, but by reading. It read billions of DNA letters from environmental microbes and practiced one task, predicting the next letter. To succeed, it had to learn how genes are built. That gives it useful clues about what an unknown gene may do, even if no lab has ever studied it.
We evaluated LOAM across several biological benchmarks to test what it had learned from DNA alone. Across these tests, LOAM outperformed other models of a similar size and approached the performance of much larger state-of-the-art genomic models.
On one benchmark for gene function, the result was even stronger. The task was to classify genes into 128 experimentally defined enzyme-function categories using only their DNA sequence.
#1on the gene-function benchmark
Our largest LOAM model ranked first among every model we tested, including models ~10× larger.
For discovery, that matters because most genes in environmental genomes have never been experimentally characterised. Benchmarks like these show that LOAM has learned biological information that can help us rank which unknown genes are worth a closer look.
The model design, training and full test results are in our companion technical blog and paper.
Read: Introducing LOAM: Learning Biology from Long-Read Soil MetagenomesHow it fits together

For a specific soil-borne pathogen, the steps look like this. First, field data shows which microbes keep turning up where the disease is held back. Next, we sequence those microbes and piece their DNA back together into genomes. Then LOAM and target-specific tools rank candidate peptides, proteins, enzymes or other functions for testing in a suitable lab assay.
The aim is not to replace lab screening. It is to send the lab a smaller, better-informed shortlist.
Lab results then show what actually works, and those results improve the next round of ranking. Over several rounds, a discovery programme moves through three questions:
- Where in nature should we look?
- Which candidates should we test?
- How can the best ones be improved to meet the product goal?
Why Soilytix?
Soilytix is building a crop-protection discovery platform that brings together field data, DNA sequencing and AI models.
Our field data shows where potentially useful biology may be concentrated. Our own lab lets us return to those soils and generate deep long-read genomic data tailored to a specific discovery question, instead of relying only on DNA already sitting in public databases. LOAM and complementary computational tools then help search that sequence space and rank the genes, proteins and other candidates most worth testing.
Doing this well takes more than one dataset or one model. The Soilytix team combines soil ecology, molecular biology, genomics, biomarker discovery, machine learning and software engineering, together with experience turning complex biological data into working products.
At Soilytix, we use the soil as a source of new biology, starting in agriculture and aiming, over time, for the health of people and the planet.
About the authors

Tim Rajakumar
Co-founder, Chief Scientific OfficerTim co-founded Soilytix and sets its research direction, owning the science behind its assays and claims. He leads the company's applied science and co-leads its AI work on bioassets with Quentin.

Quentin Ferry
Co-founder, Chief AI ScientistQuentin co-founded Soilytix and leads its AI research for bioassets. That work includes LOAM, the family of genomic language models introduced in this post.
