Tuesday, October 6, 2026

From Genomics to AI Biology (Series 4)

From Prediction to Experiment

Can AI Become a Scientist?


By Li Lei on Oct 6, 2026

For most of the history of AI in biology, the relationship has been relatively simple:

We give AI biological data.
AI gives us a prediction.

A protein structure.
A candidate gene.
A pathogenic variant.
A cell state.
A molecular property.
A ranked list of hypotheses.

Then a scientist decides what to do next.

But that relationship is beginning to change.

AI systems are moving beyond prediction. They can search the literature, analyze data, generate hypotheses, critique ideas, use scientific software, propose experiments, interpret results, and revise their plans in response to new evidence.

In other words, AI is beginning to enter the scientific loop itself.

That raises a deeper question than whether AI can predict biology:

Can AI become a scientist?

My answer is both yes and no.

AI is increasingly capable of performing pieces of scientific work that once required human researchers. But science is not simply a collection of tasks.

Science is an iterative process of deciding:

What do we not know?

What evidence would reduce that uncertainty?

What experiment should we perform next?

What does the result mean?

And perhaps most importantly:

Which question is worth asking in the first place?

The next frontier of AI biology may therefore not be better prediction.

It may be closed-loop discovery.


Science is a loop, not a prediction

A conventional machine-learning system often looks like this:

Data → Model → Prediction

But scientific discovery looks more like this:

Observation → Hypothesis → Experiment → Result → Interpretation → New Hypothesis

And then the cycle repeats.

That difference is fundamental.

Suppose an AI model predicts that a particular gene may contribute to drought tolerance.

Useful.

But a scientist immediately faces another question:

What should I do next?

Knock out the gene?

Overexpress it?

Measure its expression in roots?

Examine chromatin accessibility?

Test multiple genetic backgrounds?

Apply drought at different developmental stages?

Measure survival, water-use efficiency, or yield?

Each experiment answers a different question.

Prediction identifies a possibility.

Experimental design determines what we learn from it.

That is why the transition from prediction to experiment matters so much.


A model predicts. An agent acts.

It helps to separate three levels of scientific AI.

A model asks:

What do I predict?

An agent asks:

What should I do next?

A scientist asks something harder:

What experiment would most reduce our uncertainty about the biological system?

These questions are related, but they are not equivalent.

A predictive model typically maps an input to an output.

An agent can potentially:

observe a goal,

form a plan,

retrieve information,

select tools,

execute analyses,

inspect results,

revise the plan,

and act again.

The process becomes a loop:

Goal → Plan → Act → Observe → Learn → Act again

That begins to resemble scientific work.

And this is no longer purely hypothetical.


From one AI assistant to a Virtual Lab

One of the most interesting examples comes from James Zou and colleagues.

In 2025, they introduced the Virtual Lab, an AI–human research framework in which multiple AI agents take on different scientific roles and collaborate as a research team [1].

Rather than asking one language model a question, the Virtual Lab includes an AI principal investigator coordinating specialized scientist agents, while a human researcher provides high-level guidance.

The team applied this framework to the design of SARS-CoV-2 nanobodies.

The agents discussed scientific strategies, challenged one another’s ideas, combined computational tools, and ultimately proposed 92 new nanobody designs [1].

Then came the critical step:

the designs were experimentally tested.

Several showed functional binding, including candidates with improved binding profiles against newer viral variants [1].

That distinction matters.

The system did not merely generate scientifically plausible language.

Its ideas encountered reality.

This is much closer to the standard I argued for in Essay 3:

prediction → hypothesis → experiment → evidence

The Virtual Lab also points toward another possibility.

Perhaps the future AI scientist will not be one giant model that knows everything.

Perhaps scientific AI will look more like a team:

one agent searches the literature,

another specializes in molecular modeling,

another critiques hypotheses,

another analyzes data,

while human scientists guide the research direction and evaluate the consequences.

Science has always been collaborative.

Scientific AI may become collaborative too.


AI as a computational scientist

Another example from Zou’s group moves in a different direction.

In 2026, Alber and colleagues introduced CellVoyager, an AI computational-biology agent designed to autonomously analyze single-cell RNA-seq data [2].

Single-cell datasets illustrate why scientific agents may become useful.

A single experiment can support an enormous number of possible analyses:

Which cell populations differ?

Which pathways change?

Which cell–cell interactions matter?

Which genes define a state?

Which trajectories exist?

Which comparison should be performed next?

The challenge is no longer simply computation.

It is the size of the hypothesis space.

CellVoyager can generate analysis ideas, write and execute code, inspect its results, and continue exploring the dataset [2].

The authors evaluated it across 76 published single-cell studies and several deeper case studies. Domain experts judged some of its generated findings to be both scientifically reasonable and creative [2].

This represents another transition:

AI as tool → AI as analyst → AI as explorer

But an important distinction remains.

CellVoyager operates mainly in the computational world.

The dataset already exists.

Someone has already collected the cells.

Someone has already performed the experiment.

The next step is harder:

Can AI decide what data should be generated next?


The most likely experiment is not always the best experiment

Imagine that we have 10,000 genes we could perturb.

A prediction model might rank them according to the probability that each affects a phenotype.

Suppose:

Gene A: 90% probability of affecting the phenotype

Gene B: 50% probability

If our only goal is finding a positive result, Gene A seems like the obvious choice.

But suppose Gene B sits at the boundary between two competing biological mechanisms.

If Gene B produces one outcome, Mechanism 1 becomes much more likely.

If it produces another, Mechanism 2 becomes more likely.

Testing Gene B may therefore teach us more.

This distinction is fundamental.

The most likely successful experiment is not necessarily the most informative experiment.

This is where ideas such as active learning, Bayesian optimization, and optimal experimental design become especially interesting for biology.

Instead of asking only:

What do I predict will happen?

the system can ask:

What experiment would most reduce my uncertainty?

That is a different form of intelligence.

And in science, it may be more valuable.


Science is an allocation problem

Experiments are not free.

They consume:

time,

money,

reagents,

samples,

instrument capacity,

field space,

animals,

and human attention.

Some experiments take hours.

Others take months.

Some biological materials are abundant.

Others are irreplaceable.

We cannot test every hypothesis.

Scientific research is therefore partly an allocation problem under uncertainty.

Given limited resources and an enormous hypothesis space:

Which experiment should we perform next?

This may eventually become one of AI’s most valuable contributions to science.

The best scientific AI may not be the system with the highest predictive accuracy.

It may be the system that helps us spend our next experiment most intelligently.

Imagine an AI saying:

I currently consider three mechanisms plausible.

Experiment A is highly likely to produce a positive result, but all three mechanisms predict that result.

Experiment B is less likely to succeed, but the competing mechanisms predict very different outcomes.

Therefore, Experiment B would provide substantially more information.

That begins to sound less like prediction.

It begins to sound like experimental reasoning.


From static papers to active scientific knowledge

There is another development from Zou and colleagues that changes a different part of science.

Scientific knowledge has traditionally been stored in static objects:

papers.

We read them.

Interpret them.

Download their code.

Try to reproduce their methods.

Then, perhaps, apply those methods elsewhere.

In 2026, Miao and colleagues introduced Paper2Agent, a system that transforms a scientific paper into an interactive AI agent [3].

Instead of representing a paper only as text, Paper2Agent creates something closer to a virtual corresponding author.

The resulting agent can answer detailed questions about the work, apply its methods to new data, and interact with agents generated from other papers [3].

This is conceptually fascinating.

The scientific literature may begin to move from:

static knowledge

toward

executable knowledge.

Imagine a future in which a genomic-methods paper is not merely something you read.

Its agent can:

explain the assumptions,

run the method,

apply it to your dataset,

show where it may fail,

and interact with another agent representing a complementary approach.

The literature itself could become part of the discovery system rather than simply its archive.


From information assistant to closed-loop learner

Taken together, these developments suggest a progression in scientific AI.

Stage 1: AI as information assistant

“Find and summarize what is known.”

Stage 2: AI as predictor

“Given this data, what is likely to happen?”

Stage 3: AI as analyst

“Analyze these data and identify interesting patterns.”

Stage 4: AI as collaborator

“Generate and critique hypotheses with me.”

Stage 5: AI as experiment designer

“What experiment should we perform next?”

Stage 6: AI as closed-loop learner

“Perform the experiment, observe the outcome, update your model, and choose the next experiment.”

We are moving through these stages remarkably quickly.

The final stage changes the relationship between AI and biology fundamentally.

Because now AI is not merely learning from biological history.

It is helping create new biological experience.


Closing the loop with the physical world

This is the idea behind a self-driving laboratory.

A simplified loop looks like this:

Propose experiment

↓

Execute experiment

↓

Measure result

↓

Update model

↓

Select next experiment

↓

Repeat

Self-driving laboratories combine machine learning, robotics, automated measurement, and experimental design so that this loop can repeat with progressively less manual intervention.

A 2026 review in Nature Reviews Chemistry describes the field’s movement from automation of individual laboratory tasks toward integrated systems capable of proposing, executing, and interpreting experiments [4].

But the review also emphasizes several important challenges:

scalability

generalizability

and

provenance-complete experimentation [4].

That last requirement deserves particular attention.

If an autonomous system performs hundreds or thousands of experiments, we need a complete record of:

what it did,

why it did it,

which model version made the decision,

which parameters were used,

what measurements were produced,

how those measurements were processed,

and why the next experiment was selected.

Automation without provenance could generate more experiments while making the resulting science less understandable.

The goal should not simply be autonomous experimentation.

It should be reproducible autonomous experimentation.


Biology is beginning to close the loop

Living systems make closed-loop experimentation much harder than many engineering problems.

Cells adapt.

Organisms develop.

Phenotypes depend on history.

Genotype interacts with environment.

Measurements are noisy.

The same intervention can produce different outcomes depending on context.

But early closed-loop biological systems are beginning to appear.

A 2026 bioRxiv preprint described a closed-loop robot scientist capable of applying multiple types of interventions to living biological systems while collecting imaging data [5].

In one demonstration, the system used active learning to select later interventions based on uncertainty from earlier observations.

Another 2026 bioRxiv preprint described autonomous agents designing protein variants, sending them through robotic construction and characterization, learning from experimental feedback, and selecting subsequent variants to explore [6].

The system operated iteratively across protein sequence space and reported enzymes with altered substrate specificity [6].

Both studies are preprints, so their conclusions should be treated as emerging evidence rather than established findings.

But the conceptual shift is important:

AI is beginning to learn by acting on biological systems and observing the consequences.


Learning from data versus learning from experience

Most biological AI currently learns this way:

Past observations → Model

Humans perform experiments.

Humans build datasets.

AI learns from the accumulated record.

Closed-loop systems introduce another possibility:

Model → Experiment → Observation → Model update

Now AI participates in creating its own future training data.

That difference may prove profound.

Instead of asking:

What patterns exist in the data we already collected?

the system can ask:

What data do I need in order to answer the question?

This is much closer to scientific learning.

And it could be particularly powerful in biology, where high-quality labeled data are often scarce.

The answer to data scarcity may not always be:

generate a larger dataset.

It may be:

generate a more informative dataset.


Toward generalist biological AI

These developments connect to a broader vision.

A 2026 Nature Biotechnology review co-authored by James Zou and other leaders in biological AI describes the emergence of generalist biological AI: systems capable of integrating information across DNA, RNA, proteins, and cellular systems while connecting specialized models, agents, experimental validation, and virtual-cell-like simulations [7].

That vision is very different from building one model for one task.

Real biology is connected across scales.

A DNA variant can alter regulatory activity.

Regulatory activity changes expression.

Expression changes cellular state.

Cellular state changes phenotype.

Phenotype depends on environment.

Scientific AI will eventually need to connect these levels as well.

Agents may become the orchestration layer through which specialized models, datasets, scientific literature, and experimental tools work together.

But connecting tools is still not the same as doing science.

A harder problem remains.


An agent that performs experiments is not automatically a scientist

It is tempting to look at these advances and declare:

The AI scientist has arrived.

I think that conclusion is premature.

AI may increasingly perform:

literature search,

data analysis,

coding,

hypothesis generation,

candidate prioritization,

experimental planning,

instrument control,

and model updating.

But scientific research also contains another class of decisions.

Which problem deserves attention?

Why is this question important?

Which assumptions should we challenge?

Which anomalous result should we pursue rather than discard?

When should we abandon the original hypothesis?

What level of evidence is sufficient?

What risks are acceptable?

What discovery would actually matter?

These are not merely prediction tasks.

And they are not easily captured by a single objective function.

AI can optimize an objective.

But someone still has to decide:

What should science optimize?

That may remain one of the deepest human responsibilities.


When AI touches biology, errors become different

The move from analysis to action also changes the nature of error.

If an LLM incorrectly summarizes a paper, that is an informational failure.

If an agent chooses the wrong statistical test, it can distort a scientific conclusion.

If an agent selects an inappropriate experiment, resources are wasted.

If an autonomous system controls physical interventions, mistakes can affect:

cells,

organisms,

chemical systems,

equipment,

and potentially people.

Scientific autonomy should therefore not simply be measured by:

How many human steps can we remove?

A mature scientific agent also needs:

uncertainty estimation,

permission boundaries,

traceability,

provenance,

reproducibility,

fail-safe behavior,

and clear points for human escalation.

The closer AI moves toward the physical world, the more important these capabilities become.

Not because we should prevent scientific automation.

Because trustworthy autonomy requires constraints.


The "Move 37" problem in science

But there is an interesting tension.

Humans should supervise scientific AI.

Yet human intuition cannot always be the final standard.

AlphaGo’s famous "Move 37" is a useful analogy.

Expert Go players initially found the move highly unusual.

Its value became clear only later.

Science may eventually encounter its own "Move 37".

An AI system might propose:

a gene no expert would prioritize,

a protein sequence that looks biologically strange,

a regulatory design outside conventional rules,

a drug combination nobody considered,

or an experiment that appears unlikely to succeed.

Should we reject it because experts disagree?

No.

But neither should we accept it because AI suggested it.

Science already has a mechanism for resolving that tension.

Experiment.

That may become one of AI’s most exciting roles.

AI can expand the hypothesis space beyond human intuition.

Humans can evaluate whether the hypothesis is meaningful and safe enough to test.

Nature decides whether it survives.

Perhaps biology’s "Move 37" will not be an answer produced by AI.

Perhaps it will be:

an experiment no human scientist thought to perform.


The most valuable AI may ask better questions

This brings me to what I think is the deepest point.

We often evaluate AI by the quality of its answers.

But science advances because of questions.

A more mature scientific AI might not simply answer:

Which gene causes this phenotype?

Instead, it might ask:

Which experiment would distinguish among the competing explanations for this phenotype?

That requires the system to represent not only what it believes.

It must also represent:

what it does not know,

where competing hypotheses disagree,

and

what evidence would discriminate among them.

In other words:

Scientific intelligence may require modeling ignorance as carefully as modeling knowledge.

Scientists rarely begin with certainty.

We begin with uncertainty.

And much of scientific judgment lies in deciding:

Which uncertainty is worth reducing next?


What should remain human?

I do not think the future laboratory will divide neatly into:

human scientist

versus

AI scientist.

A more plausible future is a hybrid research organization.

AI systems may increasingly handle:

large-scale literature synthesis,

routine analyses,

candidate generation,

parameter optimization,

experimental scheduling,

instrument control,

and systematic exploration.

Human scientists may spend proportionally more time on:

problem formulation,

conceptual synthesis,

causal interpretation,

unexpected observations,

standards of evidence,

risk,

ethics,

and deciding which discoveries matter.

The Virtual Lab already hints at this structure: human direction combined with multiple specialized AI agents [1].

Paper2Agent suggests that published scientific knowledge itself may eventually participate in such teams [3].

Perhaps the future research group will include:

  • human scientists
  • specialized AI agents

  • scientific models

  • interactive literature

  • automated laboratories

all connected through a shared discovery loop.


The scientist may become an architect of discovery

This brings me back to the first essay in this series.

I argued that biologists should not simply become users of AI tools.

They should become AI system thinkers.

Closed-loop discovery makes that idea even more important.

The future scientist may increasingly design not only individual experiments, but entire discovery systems.

She may define:

the biological question,

the hypothesis space,

the permitted experimental actions,

the relevant biological constraints,

the uncertainty model,

the evidence threshold,

the safety boundaries,

the stopping criteria,

and the points at which human review is mandatory.

AI can then explore within those boundaries.

Scientific expertise does not disappear.

It moves to another level.

From performing every analysis

to designing the system that performs analyses.

From manually selecting every experiment

to defining how experiments should be selected.

From reading every result

to recognizing which unexpected result deserves attention.

The scientist becomes not only an investigator.

She becomes an architect of discovery.


So, can AI become a scientist?

Parts of science clearly can be automated.

And the boundary is expanding rapidly.

AI can already participate in:

literature synthesis,

data analysis,

hypothesis generation,

scientific debate,

candidate design,

experimental prioritization,

and increasingly the loop between prediction and experiment [1–7].

But I am not convinced that the most important question is:

When will AI replace scientists?

That framing is too narrow.

The more interesting question is:

What kind of science becomes possible when humans and AI explore the unknown together?

Humans bring:

biological intuition,

problem framing,

skepticism,

context,

values,

experience,

and scientific taste.

AI brings:

scale,

search,

memory,

computation,

optimization,

and the ability to explore enormous possibility spaces systematically.

Robotics connects those capabilities to physical experimentation.

The scientific literature itself may increasingly become executable and interactive.

Together, these components could change not only the speed of science.

They could change the structure of scientific discovery.


From prediction to experiment

The first generation of AI biology largely asked:

Can we predict biology?

A second generation increasingly asks:

Can AI reason over biological evidence?

Now another question is emerging:

Can AI decide what evidence to collect next?

That transition matters.

Because science is not simply the accumulation of predictions.

Science is a conversation with nature.

We propose an explanation.

Nature answers through experiment.

We revise the explanation.

And we ask again.

The next revolution in AI biology may therefore not come from a model that simply knows more biology.

It may come from systems capable of participating in that conversation:

Observe.

Hypothesize.

Experiment.

Learn.

Revise.

Repeat.

But as AI becomes increasingly capable of participating in the cycle of discovery, one question becomes more important, not less:

Who decides what is worth discovering?

For now, I believe that remains one of the deepest responsibilities of the scientist.

And perhaps this is where the human role in AI-driven science becomes not smaller—

but more consequential.


References and Further Reading

[1] Swanson, K., Wu, W., Bulaong, N. L. et al. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature 646, 716–723 (2025). doi:10.1038/s41586-025-09442-9.

[2] Alber, S., Chen, B., Sun, E. et al. CellVoyager: AI CompBio agent generates new insights by autonomously analyzing biological data. Nature Methods 23, 749–759 (2026). doi:10.1038/s41592-026-03029-6.

[3] Miao, J., Davis, J. R., Zhang, Y., Pritchard, J. K. & Zou, J. Reimagining research papers as interactive and reliable AI agents. Nature (2026). doi:10.1038/s41586-026-11044-y.

[4] Canty, R. B. & Abolhasani, M. The past, present and future of self-driving laboratories. Nature Reviews Chemistry 10, 523–537 (2026). doi:10.1038/s41570-026-00847-2.

[5] Bielawski, K., Srinivasan, K., Gaylinn, N. et al. A Closed-Loop Robot Scientist for Autonomous Biological Discovery. bioRxiv (2026). doi:10.64898/2026.09.11.751076. Preprint.

[6] Brooks, C., Notin, P. & Romero, P. A. Learning protein function through autonomous experimental interaction. bioRxiv (2026). doi:10.64898/2026.08.14.744985. Preprint.

[7] Rao, V. M., Zhang, S., Plosky, B. S. et al. Generalist biological artificial intelligence in modeling the language of life. Nature Biotechnology 44, 918–933 (2026). doi:10.1038/s41587-026-03064-w.

[8] Ghareeb, A. E., Chang, B., Mitchener, L. et al. A multi-agent system for automating scientific discovery. Nature 655, 497–505 (2026). doi:10.1038/s41586-026-10652-y.

Note on preprints

References [5] and [6] are bioRxiv preprints and have not yet undergone peer review. I include them because closed-loop biological experimentation is developing rapidly and they illustrate important emerging directions. Their findings should therefore be interpreted as preliminary rather than established evidence.

Saturday, August 29, 2026

From Genomics to AI Biology (Series 3)

 Why AI Biology Needs Better Evaluation

By Li Lei on August 26, 2026


In AI biology, we love numbers.

Training loss.
Validation loss.
Perplexity.
AUROC.
AUPRC.
Correlation.
Benchmark rankings.

A new foundation model improves one of these numbers, and we often call it progress.

Sometimes it is.

But increasingly, I find myself asking a harder question:

What does better model performance actually mean for biology?

Imagine that a DNA foundation model is trained on billions or trillions of nucleotides.

Its training loss falls.

Its validation loss falls.

Its sequence likelihood improves.

The optimization worked.

But what exactly have we learned?

Does the model understand gene regulation better?

Can it predict the effect of a variant it has never encountered?

Does its representation transfer to another species?

Can it predict what happens after a genetic perturbation?

Can it identify a biologically meaningful candidate that a simpler model would miss?

And ultimately:

Would its prediction change what experiment a scientist chooses to do next?

Those are very different questions.

Recent evaluations of DNA and single-cell foundation models make this distinction increasingly important. Large pretrained models do not automatically outperform specialized models or even simple baselines across every biological task. Performance can depend strongly on training objective, architecture, pretraining diversity, biological context, task definition, and how evaluation data are separated from training data [1–8].

At the same time, models such as AlphaGenome and Evo 2 demonstrate that large biological models can achieve remarkable capabilities when model design and evaluation are tightly connected to meaningful biological questions [9,10].

So the lesson is not:

Foundation models do not work.

Nor is it:

Bigger models will solve biology.

The more interesting lesson is:

There is no single metric that tells us whether a foundation model understands biology.

In the first essay of this series, I argued that biologists need to move beyond simply using AI tools and learn to think in AI systems.

In the second, I argued that biology is not just data — biological meaning depends on context.

Now comes the next question:

How do we know when an AI result deserves to influence science?

A lower loss is not biological understanding.

A benchmark score is not biological truth.

A prediction is not a mechanism.

And a prediction is certainly not yet a discovery.


1. When does a prediction become a discovery?

I find it useful to think about AI biology evaluation as a six-level stack.

Figure 1. When does an AI prediction become a discovery? 

AI biology evaluation can move from foundation-model health through task performance, generalization, biological validity, experimental validation, and finally scientific decision value.

At the bottom, we ask a very technical question:

Did the model learn its training objective?

At the top, we ask a much more scientific question:

Did the model help us make a better decision or discover something new?

Those are not the same thing.

A model can perform beautifully at Level 0 and disappoint at Level 1.

It can perform well on a benchmark and collapse when confronted with a new species or experimental condition.

It can generalize statistically while learning a shortcut rather than a meaningful biological relationship.

And it can generate a biologically plausible hypothesis that fails when tested experimentally.

That is why I think we need to stop talking about AI evaluation as if it were one number.

It is a chain of evidence.


2. What does training loss actually tell us?

Before asking whether a foundation model understands biology, we should ask:

Did training work?

During pretraining, researchers may monitor:

  • training loss

  • validation loss

  • cross-entropy

  • token or sequence likelihood

  • masked-token reconstruction

  • calibration

  • convergence

  • scaling behavior

These metrics are essential.

If training loss does not decrease, optimization may be failing.

If training loss continues improving while validation loss deteriorates, the model may be overfitting.

But there is a crucial distinction.

Figure 2. Lower loss does not automatically mean better biology.
Training metrics tell us whether a model is learning its optimization objective. They do not by themselves establish downstream biological capability.


A lower validation loss tells us:

the model has become better at its mathematical objective.

It does not automatically tell us:

the model understands enhancers better;

predicts gene expression better;

generalizes across species better;

predicts perturbations better;

or makes better scientific decisions.

This distinction is particularly important for foundation models.

A masked DNA language model, an autoregressive nucleotide model, and a model using different tokenization schemes may optimize very different objectives. Their losses may not even be directly comparable.

Recent benchmarking supports this caution.

Feng and colleagues evaluated several DNA foundation models across genomic and genetic applications and found substantial task-dependent differences. General-purpose models were competitive for some problems, such as pathogenic-variant identification, while performing less strongly on others, including gene-expression prediction and causal-QTL identification [1].

Tang and colleagues similarly found that pretrained genomic-language-model representations did not consistently outperform conventional sequence representations or strong supervised models on biologically aligned regulatory-genomics tasks [2].

The important lesson is:

A foundation model can become a better model of its training data without necessarily becoming a better model of the biology we care about.


3. Foundation-model evaluation is a portfolio, not one score

Instead of searching for one universal metric, I think biological foundation models need several complementary types of evaluation.

Figure 3. Foundation-model evaluation is a portfolio, not one score.
Intrinsic, downstream, generalization, and biological evaluation answer fundamentally different questions.

Intrinsic evaluation

Did the model learn its pretraining objective?

Loss, likelihood, masked-token recovery and convergence belong here.

Downstream evaluation

Can the learned representation solve biologically meaningful tasks?

For DNA models, this might include:

variant-effect prediction,

enhancer classification,

gene-expression prediction,

splice prediction,

or regulatory activity.

For single-cell models:

cell-state representation,

perturbation response,

cell-type prediction,

or gene-expression reconstruction.

Generalization evaluation

Does the model still work when it encounters biology that is meaningfully different from what it saw during training?

New species?

New populations?

New tissues?

New sequence families?

New perturbations?

Biological evaluation

Do the predictions survive independent biological evidence?

Genetics?

Evolution?

Functional assays?

Perturbation?

Prospective validation?

These categories answer different questions.

That is why statements such as:

“Model A has lower perplexity, therefore it is better for biology”

are too strong.

Scale can create new capability.

But evaluation tells us what capability was actually created.


4. Zero-shot performance is useful — but it needs careful interpretation

One particularly interesting promise of foundation models is zero-shot prediction.

Imagine introducing a nucleotide variant into a sequence.

A genomic language model can compare the probability or likelihood assigned to the reference and alternative sequence.

If the variant disrupts a strongly learned sequence pattern, the model may assign the altered sequence a lower probability.

That creates a variant-effect score without training directly on pathogenicity labels.

Conceptually, this is powerful.

But recent work shows why evaluation matters.

Alfisi and colleagues systematically compared DNA foundation models for zero-shot variant-effect prediction and found that performance depended on more than model size. Architecture, training objective, sequence context, and diversity of pretraining data all mattered [3].

One particularly interesting observation was the value of multi-species training.

Evolution has already performed billions of years of natural experiments.

Training across diverse species may therefore expose models to patterns of biological constraint that are difficult to learn from one genome alone.

This suggests something deeper:

the composition of the training data may matter as much as the number of model parameters.

And it reinforces why foundation-model evaluation cannot be reduced to a scaling curve.


5. Strong baselines keep us scientifically honest

One of the most useful lessons from recent AI biology studies is remarkably simple:

Always compare a sophisticated model against a strong simple baseline.

This has become especially clear in single-cell biology.

Kedzierska and colleagues evaluated Geneformer and scGPT in zero-shot settings and found that their representations did not consistently outperform simpler established approaches [5].

Ahlmann-Eltze and colleagues examined prediction of transcriptional responses to genetic perturbation and found that several deep-learning approaches did not consistently outperform deliberately simple linear baselines [6].

This does not mean that single-cell foundation models have failed.

It means that a complicated model should earn its complexity.

If a very large model cannot outperform a linear model or strong task-specific baseline on the capability being claimed, that tells us something important.

Perhaps the objective needs improvement.

Perhaps the benchmark is too easy.

Perhaps the training data do not contain enough causal information.

Perhaps the representation captures cell identity well but not intervention response.

Emerging bioRxiv work is pushing this idea further. The scFME framework, for example, evaluates single-cell foundation models specifically for in-silico perturbation, comparing predicted changes against control, zero, and random perturbation baselines rather than simply asking whether embeddings appear structured [7].

That distinction matters.

A beautiful embedding is not necessarily a useful perturbation model.

A beautiful UMAP is not a discovery.


6. Is the model generalizing — or recognizing something familiar?

This is one of the most important problems in biological machine learning.

Biological observations are rarely independent.

DNA sequences share homology.

Proteins belong to families.

Individuals share ancestry.

Cells may come from the same donor.

Species share evolutionary history.

Experiments share protocols and batch effects.

A random train-test split can therefore create the illusion of generalization.

Figure 4. Unseen rows do not necessarily mean unseen biology.
A random split may place highly homologous biological examples in both training and test sets. Biology-aware splits provide a stronger test of genuine generalization.

Suppose Gene A1 and Gene A2 are in training.

Gene A3 is placed in the test set.

Technically, A3 is unseen.

But if A3 is highly homologous to A1 and A2, has the model really encountered new biology?

Not necessarily.

Recent bioRxiv work by Rafi and colleagues examined this problem explicitly, demonstrating how homology can create leakage between training and evaluation partitions in genome-trained sequence models [4].

This leads to one of the most important principles in biological AI evaluation:

A test example is not truly “unseen” simply because its row was absent from the training set.

Biological independence matters.

Depending on the scientific question, stronger evaluation may require holding out:

an entire chromosome,

a gene family,

a homology cluster,

a donor,

a population,

an environmental condition,

or an entire species.

These evaluations are harder.

But discovery happens precisely when models encounter something they have not already seen.


7. Did the model learn biology — or a shortcut?

Even if a model generalizes statistically, we still need to ask:

What did it actually learn?

Suppose an AI system predicts genes associated with drought tolerance.

Perhaps the top candidates are:

expressed in the relevant tissue,

responsive under drought,

near genetic associations,

supported by chromatin accessibility,

connected to stress-response pathways,

or evolutionarily conserved.

Our confidence increases.

But perhaps the model instead learned:

GC content,

gene length,

annotation density,

batch identity,

phylogenetic similarity,

or another feature correlated with the label.

Both models might produce high AUROC.

Only one may teach us useful biology.

This is where interpretability becomes valuable.

Not because every neural network must produce a simple human-readable explanation, but because probing model representations can reveal whether recognizable biological features have emerged.

Evo 2, for example, reported internal features associated with exon-intron boundaries, transcription-factor-binding sites, protein structural properties, and prophage-related sequence [10].

That is encouraging.

But interpretability is still not the final test.

A model feature that resembles biology is evidence.

It is not yet mechanism.


8. AlphaGenome: evaluation closer to biological measurement

AlphaGenome offers an interesting example of evaluation that is more closely aligned with biological observables.

Rather than producing only a generic embedding, AlphaGenome predicts thousands of genomic tracks from long DNA sequences, including:

gene expression,

transcription initiation,

chromatin accessibility,

histone modifications,

transcription-factor binding,

chromatin contacts,

and splicing [9].

Its variant-effect evaluations then ask how sequence changes alter these measurable molecular outputs.

That is scientifically attractive because the evaluation moves closer to quantities biologists actually measure.

The principle is broader than AlphaGenome:

The closer an AI evaluation is to a meaningful biological measurement or intervention, the stronger the scientific claim we can make.

Predicting a benchmark label is useful.

Predicting an experimentally measurable consequence is stronger.

Predicting a new experimental outcome prospectively is stronger still.


9. Can nature prove the model wrong?

For me, this is where AI biology becomes most exciting.

A good scientific prediction should eventually confront reality.

Imagine an AI system predicts that a regulatory sequence controls a drought-responsive gene.

What would test that hypothesis?

CRISPR perturbation?

Reporter assay?

Perturb-seq?

Expression profiling?

Chromatin accessibility?

Protein binding?

A field experiment?

The exact experiment depends on the biological question.

But the principle is universal:

A scientific AI system should ideally generate hypotheses that can be falsified.

We should be able to ask:

What result would prove this prediction wrong?

Loss cannot answer that.

AUROC cannot answer that.

A benchmark cannot completely answer that.

Eventually, nature has to answer.

That is why prospective validation is so powerful.

Make the prediction first.

Then collect new evidence.

Blind prediction challenges such as CASP became influential in structural biology because predictions were evaluated against experimentally determined structures that participants did not know when making their predictions.

The answer existed outside the model.

AI biology needs more tests with this character:

prospective prediction,

blind validation,

independent experiments,

orthogonal assays,

and perturbational testing.

Instead of asking only:

Can the model explain biology we already measured?

we should increasingly ask:

Can it correctly predict biology we have not measured yet?

That is much closer to discovery.


10. Does AI change the next scientific decision?

There is one final level of evaluation that I think receives too little attention.

Does the AI system improve a decision?

Scientists rarely build models because they simply want more predictions.

Eventually, we need to decide:

Which gene should we perturb?

Which variant should we prioritize?

Which molecule should we synthesize?

Which experiment should we run?

Which hypothesis deserves another six months of work?

Imagine two models ranking 10,000 genes.

Model A:

AUROC = 0.92

Model B:

AUROC = 0.89

Model A appears better.

But the laboratory can test only ten genes.

Among those ten predictions:

Model A contains two validated candidates.

Model B contains six.

Which model is more scientifically useful?

Probably Model B.

This suggests that AI biology may need to pay more attention to metrics such as:

top-k enrichment,

experimental hit rate,

number needed to test,

cost per validated candidate,

search-space reduction,

and information gained from the next experiment.

These metrics move us from prediction performance toward decision value.

And ultimately, decision value may be one of the most important measures of AI's contribution to science.


11. What about scientific AI agents?

Evaluation becomes even more difficult when we move from models to agents.

A model might perform one task:

predict a variant effect.

An AI agent might:

interpret a scientific question,

search literature,

retrieve data,

write analysis code,

choose statistical methods,

run tools,

interpret results,

generate hypotheses,

and recommend an experiment.

How do we evaluate such a system?

Only the final answer?

The retrieved evidence?

The analysis code?

The statistical choices?

The reproducibility?

The proposed experiment?

One subtle error early in the workflow can propagate through everything downstream.

ScienceAgentBench was developed specifically to test data-driven scientific agents using 102 tasks derived from 44 peer-reviewed publications [12].

Even advanced agent configurations solved only a fraction of the tasks reliably.

That does not mean scientific agents are unpromising.

It means we should be cautious about moving from impressive demonstrations to claims of autonomous scientific discovery.

The more autonomous an AI system becomes,

the more rigorous its evaluation must become.

Not less.


Human agreement is not the final benchmark

There is another complication.

Human expertise is essential.

But human agreement cannot always be the ultimate standard.

AlphaGo's famous Move 37 is a useful reminder.

The move initially appeared highly unusual to expert players.

Its value became clear later.

Science may eventually encounter its own versions of Move 37.

An AI system may propose:

an unexpected gene,

a surprising regulatory mechanism,

an unusual protein sequence,

a non-obvious molecular interaction,

or an experiment an experienced scientist would never have prioritized.

If our evaluation asks only:

Does the expert agree with the AI?

we may reject precisely the discoveries AI could help us make.

So human judgment matters.

But the stronger question is:

Does reality agree?

Does the molecule function?

Does the sequence regulate expression?

Does the protein fold?

Does the perturbation change the phenotype?

Does the result replicate?

The ultimate judge of scientific AI should not be AI itself.

And it should not always be human intuition.

It should be evidence.


From prediction to discovery

We are entering an extraordinary period in biology.

Models such as AlphaGenome and Evo 2 demonstrate capabilities that would have seemed extraordinarily ambitious only a few years ago [9,10].

Foundation models are learning increasingly rich representations of DNA, proteins, cells, and molecules.

Scientific AI agents are beginning to interact with multi-step research workflows.

The opportunity is real.

But greater capability creates a greater need for scientific discipline.

As models become more powerful, generating plausible predictions becomes easier.

Determining which predictions deserve our trust becomes harder.

So perhaps the most important question in AI biology is no longer simply:

Can AI make a prediction?

We should ask:

Did the model actually learn?

Then:

Does it work on the biological task?

Then:

Does it generalize to genuinely new biology?

Then:

Did it learn meaningful biology rather than a shortcut?

Then:

Can we test its prediction?

Then:

Did it improve the next scientific decision?

And eventually:

Did it help us discover something we did not know before?

That is the journey from prediction to discovery.

Training loss tells us whether optimization is working.

Benchmarks tell us whether a model performs on defined tasks.

Strong baselines tell us whether complexity adds value.

Generalization tests tell us whether the capability survives unfamiliar biology.

Biological validation tells us whether the prediction survives context.

Experiment tells us whether nature agrees.

Decision value tells us whether AI actually changes science.

None of these alone is enough.

Together, they form a much stronger definition of progress.

The future of AI biology will therefore not be defined only by who builds the largest model,

trains on the most data,

achieves the lowest loss,

or reaches the top of a leaderboard.

It will be defined by our ability to distinguish between:

a well-trained model,

a strong benchmark result,

a biologically meaningful prediction,

a plausible hypothesis,

reliable evidence,

and a real discovery.

Those are not the same thing.

Knowing the difference is scientific judgment.

And as AI becomes more powerful, I believe that judgment will become more valuable — not less.


References and Further Reading

[1] Feng, H., Wu, L., Zhao, B. et al. Benchmarking DNA foundation models for genomic and genetic tasks. Nature Communications 16, 10780 (2025). doi:10.1038/s41467-025-65823-8.

[2] Tang, Z., Somia, N., Yu, Y. & Koo, P. K. Evaluating the representational power of pre-trained DNA language models for regulatory genomics. Genome Biology 26, 203 (2025). doi:10.1186/s13059-025-03674-8.

[3] Alfisi, I., Ciapi, F., Baragli, M. & Magi, A. Benchmarking DNA foundation models for zero-shot variant effect prediction shows the importance of context, training, and architecture. Genome Biology (2026). doi:10.1186/s13059-026-04238-0.

[4] Rafi, A. M., Kiyota, B., Yachie, N. & de Boer, C. Detecting and avoiding homology-based data leakage in genome-trained sequence models. bioRxiv (2025). doi:10.1101/2025.01.22.634321. Preprint.

[5] Kedzierska, K. Z., Crawford, L., Amini, A. P. & Lu, A. X. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biology 26, 101 (2025). doi:10.1186/s13059-025-03574-x.

[6] Ahlmann-Eltze, C., Huber, W. & Anders, S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nature Methods 22, 1657–1661 (2025). doi:10.1038/s41592-025-02772-6.

[7] Boylan, J., Solovyeva, E., Bouiller, T. et al. Single Cell Foundation Models Evaluation (scFME) for In-Silico Perturbation. bioRxiv (2025). doi:10.1101/2025.09.22.677811. Preprint.

[8] Weidener, L. S., Brkić, M., Jovanović, M., Ulgac, E. & Meduri, A. VCBench: A Multi-Dimensional Benchmark for Single-Cell Foundation Models. bioRxiv (2026). doi:10.64898/2026.06.18.733146. Preprint.

[9] Avsec, Ž., Latysheva, N., Cheng, J. et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 649, 1206–1218 (2026). doi:10.1038/s41586-025-10014-0.

[10] Brixi, G., Durrant, M. G., Ku, J. et al. Genome modelling and design across all domains of life with Evo 2. Nature 652, 1349–1361 (2026). doi:10.1038/s41586-026-10176-5.

[11] Kapoor, S., Cantrell, E. M., Peng, K. et al. REFORMS: Consensus-based Recommendations for Machine-learning-based Science. Science Advances 10, eadk3452 (2024). doi:10.1126/sciadv.adk3452.

[12] Chen, Z., Chen, S., Ning, Y. et al. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. International Conference on Learning Representations (ICLR) (2025).

Note on preprints

Several references above are bioRxiv preprints and have not yet undergone peer review. I include them because biological foundation-model evaluation is developing rapidly and some important methodological work appears first as preprints. Their findings should therefore be interpreted as emerging evidence rather than established consensus.