Tuesday, October 6, 2026

From Genomics to AI Biology (Series 4)

From Prediction to Experiment

Can AI Become a Scientist?


By Li Lei on Oct 6, 2026

For most of the history of AI in biology, the relationship has been relatively simple:

We give AI biological data.
AI gives us a prediction.

A protein structure.
A candidate gene.
A pathogenic variant.
A cell state.
A molecular property.
A ranked list of hypotheses.

Then a scientist decides what to do next.

But that relationship is beginning to change.

AI systems are moving beyond prediction. They can search the literature, analyze data, generate hypotheses, critique ideas, use scientific software, propose experiments, interpret results, and revise their plans in response to new evidence.

In other words, AI is beginning to enter the scientific loop itself.

That raises a deeper question than whether AI can predict biology:

Can AI become a scientist?

My answer is both yes and no.

AI is increasingly capable of performing pieces of scientific work that once required human researchers. But science is not simply a collection of tasks.

Science is an iterative process of deciding:

What do we not know?

What evidence would reduce that uncertainty?

What experiment should we perform next?

What does the result mean?

And perhaps most importantly:

Which question is worth asking in the first place?

The next frontier of AI biology may therefore not be better prediction.

It may be closed-loop discovery.


Science is a loop, not a prediction

A conventional machine-learning system often looks like this:

Data → Model → Prediction

But scientific discovery looks more like this:

Observation → Hypothesis → Experiment → Result → Interpretation → New Hypothesis

And then the cycle repeats.

That difference is fundamental.

Suppose an AI model predicts that a particular gene may contribute to drought tolerance.

Useful.

But a scientist immediately faces another question:

What should I do next?

Knock out the gene?

Overexpress it?

Measure its expression in roots?

Examine chromatin accessibility?

Test multiple genetic backgrounds?

Apply drought at different developmental stages?

Measure survival, water-use efficiency, or yield?

Each experiment answers a different question.

Prediction identifies a possibility.

Experimental design determines what we learn from it.

That is why the transition from prediction to experiment matters so much.


A model predicts. An agent acts.

It helps to separate three levels of scientific AI.

A model asks:

What do I predict?

An agent asks:

What should I do next?

A scientist asks something harder:

What experiment would most reduce our uncertainty about the biological system?

These questions are related, but they are not equivalent.

A predictive model typically maps an input to an output.

An agent can potentially:

observe a goal,

form a plan,

retrieve information,

select tools,

execute analyses,

inspect results,

revise the plan,

and act again.

The process becomes a loop:

Goal → Plan → Act → Observe → Learn → Act again

That begins to resemble scientific work.

And this is no longer purely hypothetical.


From one AI assistant to a Virtual Lab

One of the most interesting examples comes from James Zou and colleagues.

In 2025, they introduced the Virtual Lab, an AI–human research framework in which multiple AI agents take on different scientific roles and collaborate as a research team [1].

Rather than asking one language model a question, the Virtual Lab includes an AI principal investigator coordinating specialized scientist agents, while a human researcher provides high-level guidance.

The team applied this framework to the design of SARS-CoV-2 nanobodies.

The agents discussed scientific strategies, challenged one another’s ideas, combined computational tools, and ultimately proposed 92 new nanobody designs [1].

Then came the critical step:

the designs were experimentally tested.

Several showed functional binding, including candidates with improved binding profiles against newer viral variants [1].

That distinction matters.

The system did not merely generate scientifically plausible language.

Its ideas encountered reality.

This is much closer to the standard I argued for in Essay 3:

prediction → hypothesis → experiment → evidence

The Virtual Lab also points toward another possibility.

Perhaps the future AI scientist will not be one giant model that knows everything.

Perhaps scientific AI will look more like a team:

one agent searches the literature,

another specializes in molecular modeling,

another critiques hypotheses,

another analyzes data,

while human scientists guide the research direction and evaluate the consequences.

Science has always been collaborative.

Scientific AI may become collaborative too.


AI as a computational scientist

Another example from Zou’s group moves in a different direction.

In 2026, Alber and colleagues introduced CellVoyager, an AI computational-biology agent designed to autonomously analyze single-cell RNA-seq data [2].

Single-cell datasets illustrate why scientific agents may become useful.

A single experiment can support an enormous number of possible analyses:

Which cell populations differ?

Which pathways change?

Which cell–cell interactions matter?

Which genes define a state?

Which trajectories exist?

Which comparison should be performed next?

The challenge is no longer simply computation.

It is the size of the hypothesis space.

CellVoyager can generate analysis ideas, write and execute code, inspect its results, and continue exploring the dataset [2].

The authors evaluated it across 76 published single-cell studies and several deeper case studies. Domain experts judged some of its generated findings to be both scientifically reasonable and creative [2].

This represents another transition:

AI as tool → AI as analyst → AI as explorer

But an important distinction remains.

CellVoyager operates mainly in the computational world.

The dataset already exists.

Someone has already collected the cells.

Someone has already performed the experiment.

The next step is harder:

Can AI decide what data should be generated next?


The most likely experiment is not always the best experiment

Imagine that we have 10,000 genes we could perturb.

A prediction model might rank them according to the probability that each affects a phenotype.

Suppose:

Gene A: 90% probability of affecting the phenotype

Gene B: 50% probability

If our only goal is finding a positive result, Gene A seems like the obvious choice.

But suppose Gene B sits at the boundary between two competing biological mechanisms.

If Gene B produces one outcome, Mechanism 1 becomes much more likely.

If it produces another, Mechanism 2 becomes more likely.

Testing Gene B may therefore teach us more.

This distinction is fundamental.

The most likely successful experiment is not necessarily the most informative experiment.

This is where ideas such as active learning, Bayesian optimization, and optimal experimental design become especially interesting for biology.

Instead of asking only:

What do I predict will happen?

the system can ask:

What experiment would most reduce my uncertainty?

That is a different form of intelligence.

And in science, it may be more valuable.


Science is an allocation problem

Experiments are not free.

They consume:

time,

money,

reagents,

samples,

instrument capacity,

field space,

animals,

and human attention.

Some experiments take hours.

Others take months.

Some biological materials are abundant.

Others are irreplaceable.

We cannot test every hypothesis.

Scientific research is therefore partly an allocation problem under uncertainty.

Given limited resources and an enormous hypothesis space:

Which experiment should we perform next?

This may eventually become one of AI’s most valuable contributions to science.

The best scientific AI may not be the system with the highest predictive accuracy.

It may be the system that helps us spend our next experiment most intelligently.

Imagine an AI saying:

I currently consider three mechanisms plausible.

Experiment A is highly likely to produce a positive result, but all three mechanisms predict that result.

Experiment B is less likely to succeed, but the competing mechanisms predict very different outcomes.

Therefore, Experiment B would provide substantially more information.

That begins to sound less like prediction.

It begins to sound like experimental reasoning.


From static papers to active scientific knowledge

There is another development from Zou and colleagues that changes a different part of science.

Scientific knowledge has traditionally been stored in static objects:

papers.

We read them.

Interpret them.

Download their code.

Try to reproduce their methods.

Then, perhaps, apply those methods elsewhere.

In 2026, Miao and colleagues introduced Paper2Agent, a system that transforms a scientific paper into an interactive AI agent [3].

Instead of representing a paper only as text, Paper2Agent creates something closer to a virtual corresponding author.

The resulting agent can answer detailed questions about the work, apply its methods to new data, and interact with agents generated from other papers [3].

This is conceptually fascinating.

The scientific literature may begin to move from:

static knowledge

toward

executable knowledge.

Imagine a future in which a genomic-methods paper is not merely something you read.

Its agent can:

explain the assumptions,

run the method,

apply it to your dataset,

show where it may fail,

and interact with another agent representing a complementary approach.

The literature itself could become part of the discovery system rather than simply its archive.


From information assistant to closed-loop learner

Taken together, these developments suggest a progression in scientific AI.

Stage 1: AI as information assistant

“Find and summarize what is known.”

Stage 2: AI as predictor

“Given this data, what is likely to happen?”

Stage 3: AI as analyst

“Analyze these data and identify interesting patterns.”

Stage 4: AI as collaborator

“Generate and critique hypotheses with me.”

Stage 5: AI as experiment designer

“What experiment should we perform next?”

Stage 6: AI as closed-loop learner

“Perform the experiment, observe the outcome, update your model, and choose the next experiment.”

We are moving through these stages remarkably quickly.

The final stage changes the relationship between AI and biology fundamentally.

Because now AI is not merely learning from biological history.

It is helping create new biological experience.


Closing the loop with the physical world

This is the idea behind a self-driving laboratory.

A simplified loop looks like this:

Propose experiment

↓

Execute experiment

↓

Measure result

↓

Update model

↓

Select next experiment

↓

Repeat

Self-driving laboratories combine machine learning, robotics, automated measurement, and experimental design so that this loop can repeat with progressively less manual intervention.

A 2026 review in Nature Reviews Chemistry describes the field’s movement from automation of individual laboratory tasks toward integrated systems capable of proposing, executing, and interpreting experiments [4].

But the review also emphasizes several important challenges:

scalability

generalizability

and

provenance-complete experimentation [4].

That last requirement deserves particular attention.

If an autonomous system performs hundreds or thousands of experiments, we need a complete record of:

what it did,

why it did it,

which model version made the decision,

which parameters were used,

what measurements were produced,

how those measurements were processed,

and why the next experiment was selected.

Automation without provenance could generate more experiments while making the resulting science less understandable.

The goal should not simply be autonomous experimentation.

It should be reproducible autonomous experimentation.


Biology is beginning to close the loop

Living systems make closed-loop experimentation much harder than many engineering problems.

Cells adapt.

Organisms develop.

Phenotypes depend on history.

Genotype interacts with environment.

Measurements are noisy.

The same intervention can produce different outcomes depending on context.

But early closed-loop biological systems are beginning to appear.

A 2026 bioRxiv preprint described a closed-loop robot scientist capable of applying multiple types of interventions to living biological systems while collecting imaging data [5].

In one demonstration, the system used active learning to select later interventions based on uncertainty from earlier observations.

Another 2026 bioRxiv preprint described autonomous agents designing protein variants, sending them through robotic construction and characterization, learning from experimental feedback, and selecting subsequent variants to explore [6].

The system operated iteratively across protein sequence space and reported enzymes with altered substrate specificity [6].

Both studies are preprints, so their conclusions should be treated as emerging evidence rather than established findings.

But the conceptual shift is important:

AI is beginning to learn by acting on biological systems and observing the consequences.


Learning from data versus learning from experience

Most biological AI currently learns this way:

Past observations → Model

Humans perform experiments.

Humans build datasets.

AI learns from the accumulated record.

Closed-loop systems introduce another possibility:

Model → Experiment → Observation → Model update

Now AI participates in creating its own future training data.

That difference may prove profound.

Instead of asking:

What patterns exist in the data we already collected?

the system can ask:

What data do I need in order to answer the question?

This is much closer to scientific learning.

And it could be particularly powerful in biology, where high-quality labeled data are often scarce.

The answer to data scarcity may not always be:

generate a larger dataset.

It may be:

generate a more informative dataset.


Toward generalist biological AI

These developments connect to a broader vision.

A 2026 Nature Biotechnology review co-authored by James Zou and other leaders in biological AI describes the emergence of generalist biological AI: systems capable of integrating information across DNA, RNA, proteins, and cellular systems while connecting specialized models, agents, experimental validation, and virtual-cell-like simulations [7].

That vision is very different from building one model for one task.

Real biology is connected across scales.

A DNA variant can alter regulatory activity.

Regulatory activity changes expression.

Expression changes cellular state.

Cellular state changes phenotype.

Phenotype depends on environment.

Scientific AI will eventually need to connect these levels as well.

Agents may become the orchestration layer through which specialized models, datasets, scientific literature, and experimental tools work together.

But connecting tools is still not the same as doing science.

A harder problem remains.


An agent that performs experiments is not automatically a scientist

It is tempting to look at these advances and declare:

The AI scientist has arrived.

I think that conclusion is premature.

AI may increasingly perform:

literature search,

data analysis,

coding,

hypothesis generation,

candidate prioritization,

experimental planning,

instrument control,

and model updating.

But scientific research also contains another class of decisions.

Which problem deserves attention?

Why is this question important?

Which assumptions should we challenge?

Which anomalous result should we pursue rather than discard?

When should we abandon the original hypothesis?

What level of evidence is sufficient?

What risks are acceptable?

What discovery would actually matter?

These are not merely prediction tasks.

And they are not easily captured by a single objective function.

AI can optimize an objective.

But someone still has to decide:

What should science optimize?

That may remain one of the deepest human responsibilities.


When AI touches biology, errors become different

The move from analysis to action also changes the nature of error.

If an LLM incorrectly summarizes a paper, that is an informational failure.

If an agent chooses the wrong statistical test, it can distort a scientific conclusion.

If an agent selects an inappropriate experiment, resources are wasted.

If an autonomous system controls physical interventions, mistakes can affect:

cells,

organisms,

chemical systems,

equipment,

and potentially people.

Scientific autonomy should therefore not simply be measured by:

How many human steps can we remove?

A mature scientific agent also needs:

uncertainty estimation,

permission boundaries,

traceability,

provenance,

reproducibility,

fail-safe behavior,

and clear points for human escalation.

The closer AI moves toward the physical world, the more important these capabilities become.

Not because we should prevent scientific automation.

Because trustworthy autonomy requires constraints.


The "Move 37" problem in science

But there is an interesting tension.

Humans should supervise scientific AI.

Yet human intuition cannot always be the final standard.

AlphaGo’s famous "Move 37" is a useful analogy.

Expert Go players initially found the move highly unusual.

Its value became clear only later.

Science may eventually encounter its own "Move 37".

An AI system might propose:

a gene no expert would prioritize,

a protein sequence that looks biologically strange,

a regulatory design outside conventional rules,

a drug combination nobody considered,

or an experiment that appears unlikely to succeed.

Should we reject it because experts disagree?

No.

But neither should we accept it because AI suggested it.

Science already has a mechanism for resolving that tension.

Experiment.

That may become one of AI’s most exciting roles.

AI can expand the hypothesis space beyond human intuition.

Humans can evaluate whether the hypothesis is meaningful and safe enough to test.

Nature decides whether it survives.

Perhaps biology’s "Move 37" will not be an answer produced by AI.

Perhaps it will be:

an experiment no human scientist thought to perform.


The most valuable AI may ask better questions

This brings me to what I think is the deepest point.

We often evaluate AI by the quality of its answers.

But science advances because of questions.

A more mature scientific AI might not simply answer:

Which gene causes this phenotype?

Instead, it might ask:

Which experiment would distinguish among the competing explanations for this phenotype?

That requires the system to represent not only what it believes.

It must also represent:

what it does not know,

where competing hypotheses disagree,

and

what evidence would discriminate among them.

In other words:

Scientific intelligence may require modeling ignorance as carefully as modeling knowledge.

Scientists rarely begin with certainty.

We begin with uncertainty.

And much of scientific judgment lies in deciding:

Which uncertainty is worth reducing next?


What should remain human?

I do not think the future laboratory will divide neatly into:

human scientist

versus

AI scientist.

A more plausible future is a hybrid research organization.

AI systems may increasingly handle:

large-scale literature synthesis,

routine analyses,

candidate generation,

parameter optimization,

experimental scheduling,

instrument control,

and systematic exploration.

Human scientists may spend proportionally more time on:

problem formulation,

conceptual synthesis,

causal interpretation,

unexpected observations,

standards of evidence,

risk,

ethics,

and deciding which discoveries matter.

The Virtual Lab already hints at this structure: human direction combined with multiple specialized AI agents [1].

Paper2Agent suggests that published scientific knowledge itself may eventually participate in such teams [3].

Perhaps the future research group will include:

  • human scientists
  • specialized AI agents

  • scientific models

  • interactive literature

  • automated laboratories

all connected through a shared discovery loop.


The scientist may become an architect of discovery

This brings me back to the first essay in this series.

I argued that biologists should not simply become users of AI tools.

They should become AI system thinkers.

Closed-loop discovery makes that idea even more important.

The future scientist may increasingly design not only individual experiments, but entire discovery systems.

She may define:

the biological question,

the hypothesis space,

the permitted experimental actions,

the relevant biological constraints,

the uncertainty model,

the evidence threshold,

the safety boundaries,

the stopping criteria,

and the points at which human review is mandatory.

AI can then explore within those boundaries.

Scientific expertise does not disappear.

It moves to another level.

From performing every analysis

to designing the system that performs analyses.

From manually selecting every experiment

to defining how experiments should be selected.

From reading every result

to recognizing which unexpected result deserves attention.

The scientist becomes not only an investigator.

She becomes an architect of discovery.


So, can AI become a scientist?

Parts of science clearly can be automated.

And the boundary is expanding rapidly.

AI can already participate in:

literature synthesis,

data analysis,

hypothesis generation,

scientific debate,

candidate design,

experimental prioritization,

and increasingly the loop between prediction and experiment [1–7].

But I am not convinced that the most important question is:

When will AI replace scientists?

That framing is too narrow.

The more interesting question is:

What kind of science becomes possible when humans and AI explore the unknown together?

Humans bring:

biological intuition,

problem framing,

skepticism,

context,

values,

experience,

and scientific taste.

AI brings:

scale,

search,

memory,

computation,

optimization,

and the ability to explore enormous possibility spaces systematically.

Robotics connects those capabilities to physical experimentation.

The scientific literature itself may increasingly become executable and interactive.

Together, these components could change not only the speed of science.

They could change the structure of scientific discovery.


From prediction to experiment

The first generation of AI biology largely asked:

Can we predict biology?

A second generation increasingly asks:

Can AI reason over biological evidence?

Now another question is emerging:

Can AI decide what evidence to collect next?

That transition matters.

Because science is not simply the accumulation of predictions.

Science is a conversation with nature.

We propose an explanation.

Nature answers through experiment.

We revise the explanation.

And we ask again.

The next revolution in AI biology may therefore not come from a model that simply knows more biology.

It may come from systems capable of participating in that conversation:

Observe.

Hypothesize.

Experiment.

Learn.

Revise.

Repeat.

But as AI becomes increasingly capable of participating in the cycle of discovery, one question becomes more important, not less:

Who decides what is worth discovering?

For now, I believe that remains one of the deepest responsibilities of the scientist.

And perhaps this is where the human role in AI-driven science becomes not smaller—

but more consequential.


References and Further Reading

[1] Swanson, K., Wu, W., Bulaong, N. L. et al. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature 646, 716–723 (2025). doi:10.1038/s41586-025-09442-9.

[2] Alber, S., Chen, B., Sun, E. et al. CellVoyager: AI CompBio agent generates new insights by autonomously analyzing biological data. Nature Methods 23, 749–759 (2026). doi:10.1038/s41592-026-03029-6.

[3] Miao, J., Davis, J. R., Zhang, Y., Pritchard, J. K. & Zou, J. Reimagining research papers as interactive and reliable AI agents. Nature (2026). doi:10.1038/s41586-026-11044-y.

[4] Canty, R. B. & Abolhasani, M. The past, present and future of self-driving laboratories. Nature Reviews Chemistry 10, 523–537 (2026). doi:10.1038/s41570-026-00847-2.

[5] Bielawski, K., Srinivasan, K., Gaylinn, N. et al. A Closed-Loop Robot Scientist for Autonomous Biological Discovery. bioRxiv (2026). doi:10.64898/2026.09.11.751076. Preprint.

[6] Brooks, C., Notin, P. & Romero, P. A. Learning protein function through autonomous experimental interaction. bioRxiv (2026). doi:10.64898/2026.08.14.744985. Preprint.

[7] Rao, V. M., Zhang, S., Plosky, B. S. et al. Generalist biological artificial intelligence in modeling the language of life. Nature Biotechnology 44, 918–933 (2026). doi:10.1038/s41587-026-03064-w.

[8] Ghareeb, A. E., Chang, B., Mitchener, L. et al. A multi-agent system for automating scientific discovery. Nature 655, 497–505 (2026). doi:10.1038/s41586-026-10652-y.

Note on preprints

References [5] and [6] are bioRxiv preprints and have not yet undergone peer review. I include them because closed-loop biological experimentation is developing rapidly and they illustrate important emerging directions. Their findings should therefore be interpreted as preliminary rather than established evidence.