Let the ground truth decide where to look

How Soilytix narrows the search for new biological crop protection

Tim Rajakumar, Quentin Ferry
1 October 2026
5 min

Crop protection is under pressure. Pathogens are becoming resistant, and regulators are taking established products off the market. The industry needs new solutions, and that means searching new sources.

Most discovery still draws on the same well-known microbes and chemistry. Soil is one of the largest sources that remains largely unexplored.

It is one of the most biologically diverse places on Earth. Its microbes have spent billions of years competing with each other and living alongside plants. Their DNA holds the recipes for the proteins, enzymes and natural compounds they use to fight, feed and survive.

Most of it has never been explored. That makes soil one of the richest untapped sources for the next generation of biological crop protection.

The real challenge: knowing where to look

Long-read sequencing now lets us read the DNA of soil microbes straight from a sample, without first growing them in a lab. Getting the DNA is no longer the hard part.

Choosing what to do with it is. A handful of soil holds thousands of species and millions of genes. Every candidate has to be tested in a lab, and lab testing is slow and expensive. You cannot test everything. The real challenge is choosing the few worth testing.

Our approach: let the field decide. Nature has already run the experiment, in real fields, under real conditions. We read the result.

Nature already ran the experiment

Picture two fields where the same crop pathogen is present.

Two fields with the same crop pathogen. Left: disease develops; the pathogen is present and the crop gets sick. Right: disease-suppressive soil; pathogens are present among other microbes, and disease stays away.

In one, the crop gets sick. In the other, the pathogen is there, but the disease never appears.

The second field may contain microbes that help keep the disease in check.

Microbes are not the only possible explanation. Crop variety, weather, management, soil chemistry, how much of the pathogen is present and other factors can all influence whether disease develops.

But across enough fields, differences in disease outcome can point to something useful. They let us ask a simple question: which microbes keep turning up where the disease stays away?

Soilytix has built a dataset spanning thousands of agricultural field samples, combining microbial community profiles with information including pathogen occurrence and observed disease outcomes.

That changes where discovery starts. Instead of screening microbes picked with little knowledge of where they came from, we can start with the ones linked to the result that matters: a healthy crop despite the pathogen.

The next question: which of their genes are worth testing?

Reading genes nobody has studied

Sequencing gives us the DNA. It does not tell us what that DNA does. That is where LOAM comes in.

Most genes in soil have never been studied, so nobody knows what they do. That is exactly where new solutions can hide, and exactly where traditional tools struggle, because they rely on comparing a gene with genes that are already known. LOAM, our family of AI models, helps us judge which of those unknown genes are worth a closer look.

LOAM learns the way people learn a language: not from a dictionary, but by reading. It read billions of DNA letters from environmental microbes and practiced one task, predicting the next letter. To succeed, it had to learn how genes are built. That gives it useful clues about what an unknown gene may do, even if no lab has ever studied it.

We evaluated LOAM across several biological benchmarks to test what it had learned from DNA alone. Across these tests, LOAM outperformed other models of a similar size and approached the performance of much larger state-of-the-art genomic models.

On one benchmark for gene function, the result was even stronger. The task was to classify genes into 128 experimentally defined enzyme-function categories using only their DNA sequence.

#1on the gene-function benchmark

Our largest LOAM model ranked first among every model we tested, including models ~10× larger.

For discovery, that matters because most genes in environmental genomes have never been experimentally characterised. Benchmarks like these show that LOAM has learned biological information that can help us rank which unknown genes are worth a closer look.

The model design, training and full test results are in our companion technical blog and paper.

Read: Introducing LOAM: Learning Biology from Long-Read Soil Metagenomes

How it fits together

Discovery pipeline in seven steps: field phenotype, ecological enrichment signal and microbial lineages (find where protective biology occurs); environmental genomes (unlock their genomic diversity); genome annotation and mining with LOAM, candidate biological actives and experimental testing (turn genomes into product candidates).
Figure: result in the field → the microbes linked to it → their DNA → LOAM and other tools rank candidates → lab testing → results sharpen the next round

For a specific soil-borne pathogen, the steps look like this. First, field data shows which microbes keep turning up where the disease is held back. Next, we sequence those microbes and piece their DNA back together into genomes. Then LOAM and target-specific tools rank candidate peptides, proteins, enzymes or other functions for testing in a suitable lab assay.

The aim is not to replace lab screening. It is to send the lab a smaller, better-informed shortlist.

Lab results then show what actually works, and those results improve the next round of ranking. Over several rounds, a discovery programme moves through three questions:

  • Where in nature should we look?
  • Which candidates should we test?
  • How can the best ones be improved to meet the product goal?

Why Soilytix?

Soilytix is building a crop-protection discovery platform that brings together field data, DNA sequencing and AI models.

Our field data shows where potentially useful biology may be concentrated. Our own lab lets us return to those soils and generate deep long-read genomic data tailored to a specific discovery question, instead of relying only on DNA already sitting in public databases. LOAM and complementary computational tools then help search that sequence space and rank the genes, proteins and other candidates most worth testing.

Doing this well takes more than one dataset or one model. The Soilytix team combines soil ecology, molecular biology, genomics, biomarker discovery, machine learning and software engineering, together with experience turning complex biological data into working products.

At Soilytix, we use the soil as a source of new biology, starting in agriculture and aiming, over time, for the health of people and the planet.

About the authors

Tim Rajakumar

Co-founder, Chief Scientific Officer

Tim co-founded Soilytix and sets its research direction, owning the science behind its assays and claims. He leads the company's applied science and co-leads its AI work on bioassets with Quentin.

Quentin Ferry

Co-founder, Chief AI Scientist

Quentin co-founded Soilytix and leads its AI research for bioassets. That work includes LOAM, the family of genomic language models introduced in this post.

Follow LOAM

One or two short notes a month. Unsubscribe any time. See our privacy notice.