AI for Science Is Moving From Prediction to Closed-Loop Research Systems
How Does AI-for-Science Find Out It's Wrong?

For decades, pharma got better at generating drug candidates.
More targets, more compounds, more screening, more automation, more spend. And yet the number of new drugs approved per inflation-adjusted billion dollars of R&D fell for most of that period. In 2012, Jack Scannell and colleagues named this pattern Eroom's law: Moore's law spelled backwards, because it described the opposite of technological progress.[1]
The inputs that scaled were ideas.
The thing that did not scale was the loop that turns an idea into validated knowledge.
That loop is simple to describe (idea, evidence, experiment, interpretation, decision) and hard to execute. A hypothesis has to be grounded in prior evidence. An experiment has to test the right thing. Results have to be interpreted correctly. A decision has to be made under uncertainty. And in biology, every step can be slow, expensive, noisy, or misleading.
This is why the latest AI-for-science wave matters. Not because AI suddenly gives science more ideas: ideas were never the bottleneck. Because software is starting to touch more of the loop itself: generating hypotheses, searching evidence, writing code, running computational experiments, designing proteins, simulating biological systems, supporting decisions.
But these loops are not equivalent. The key question is not what a system can generate. It is: how does the system find out that it is wrong?
That depends on two variables that are easy to confuse. The cost of feedback is how slow, expensive, or operationally difficult it is to get a correction from reality. The fidelity of feedback is how directly that correction reflects the thing you actually care about. The two do not move together, and that gap is the whole story. Some feedback is fast and cheap but only loosely connected to the real question. Other feedback is slow and costly yet far more faithful. The cheaper and less faithful the loop, the easier it is to scale — and the easier it is to fool yourself.

Code-closed loops: computational research agents
The easiest place to automate science is where the feedback loop is digital.
That is why Sakana AI's The AI Scientist[4][5] is important. It shows that bounded parts of the computational research workflow can be automated end-to-end:
- ideation;
- literature search;
- code writing;
- experiment execution;
- result analysis;
- manuscript drafting;
- automated review.
This is impressive. But it is also the friendliest version of the problem.
Computational research has fast feedback, clear metrics, cheap iteration, and reproducible environments. A model can run code, inspect results, change parameters, and try again.
The signal is not perfect. Bad benchmarks, leaky evaluations, weak experimental design, and automated paper generation can still create noise. But when the task is well-specified, the loop can close quickly and cheaply.
That is why code-closed research is likely to be one of the first areas where autonomous agents produce visible progress.
It is not because computation is "easy."
It is because the correction mechanism is close to the model.
Literature-closed loops: hypothesis generation
Google Co-Scientist[2][3] sits one step further away from direct reality.
It is not just a chatbot for researchers. It is a multi-agent system built around structured scientific thinking: generate hypotheses, critique them, rank them, evolve them, and refine the best candidates.

That matters because part of the hypothesis-generation loop is being formalized as software.
The old version of AI for science was answer-oriented: ask a question, get a response. The new version is search-oriented: define a problem, explore the hypothesis space, compare candidates, and propose what might be worth testing.
Inherent[9] belongs in this same category. Its thesis is more ambitious: systems that help scientists find better questions, not just answer known ones. But it is still about search over possible directions of inquiry.
This is valuable. But it is not discovery by itself.
A hypothesis can be novel, elegant, and plausible while still being wrong. A system that closes mostly against literature can improve the quality of search, but it cannot fully validate the claim.
Literature can constrain the hypothesis space.
It cannot replace reality.
Simulation-closed loops: cheap feedback, dangerous fidelity
Simulation is where the framework becomes more subtle.
A simulation can look close to software because it is cheap, fast, and computational. But epistemically, it may be far from the biological reality it is supposed to predict.
That is the trap.
Cheap feedback is not the same as faithful feedback.
CellType[11] points toward biological foundation models that simulate human biology and help prioritize what to test before expensive experimental or clinical steps. BioStack[12] points toward post-training environments built from realistic healthcare and drug discovery workflows. Insilico Medicine's longevity foundation model collaboration[13] points in a similar direction: foundation models moving into aging biology, disease-risk prediction, multimodal clinical data, and preventive medicine.
This category is important because better environments can move useful feedback earlier in the research process.
But it should not be romanticized.
A simulation is useful only if it preserves the causal structure that matters when you intervene. Otherwise, it does not reduce risk. It creates false confidence.
In software, a sandbox can be close to the real environment. In biology, the sandbox is usually a proxy. Sometimes a useful proxy. Sometimes a dangerous one.
That is why simulation is not simply "between code and biology." It is a separate case: cheap like software, but potentially unfaithful like a weak biological proxy.
The strongest AI-for-science systems will not only generate candidates. They will estimate how much trust to place in each environment.
Protein- and molecule-closed loops: biological design
The next step is where AI moves from predicting biology to designing biological objects.
Biohub[6] is one signal. Its ESM release is positioned as a world model of protein biology: not only a system to predict structures, but a substrate to map, search, and design inside the protein universe.
But the broader point is not proteins alone.
The same pattern appears in enzyme design, binder design, molecular generation, and lab-in-the-loop drug discovery. AI systems propose candidates, rank them, and decide what should move into experimental validation.
This is where the loop starts to become slower and more expensive, but also more meaningful.
A generated protein or molecule is not validated because it looks good in latent space. It is validated when it works under biological constraints: binding, stability, expression, specificity, manufacturability, toxicity, and eventually clinical relevance.
That is why the term design-make-test matters.
Genesis Molecular AI and Incyte[8] show what this looks like inside pharma: proprietary experimental data feeding foundation models across multiple drug targets, inside a design-make-test loop.
Isomorphic Labs[7] represents the therapeutic execution version of the same shift. Its $2.1B Series B is not just a funding event. It is a signal that AI-first drug design is moving from research story to capital-intensive industrial strategy.
The value of AI drug discovery will not be proven by model performance alone.
It will be proven by whether AI changes the probability, speed, or cost of producing clinically meaningful assets.
Drug discovery is not a Kaggle competition.
The model output is not the product. The molecule is not even fully the product. The product is a validated therapeutic program that survives biology, safety, manufacturing, regulation, and clinical translation.
Patient-, regulator-, and market-closed loops: therapeutic reality
The final validation loop is not computational.
It is clinical, regulatory, and commercial.
This is the part of AI-for-science that funding announcements can obscure. A model can improve molecule design, reduce experimental waste, or prioritize better candidates. But the system is not truly validated until the program survives the realities it claims to improve: animal studies, human trials, safety constraints, manufacturing, regulatory review, reimbursement, and market adoption.
This is why AI-first therapeutic companies are so capital-intensive.
The closer the loop gets to patients, the more expensive the feedback becomes. But the signal also becomes harder to fake.
A clinical endpoint is slow, noisy, and expensive. But it is not a proxy in the same way a simulation is.
That is why the category is so hard. And why it is so valuable.
Workflow-closed loops: pharma decision intelligence
Not every valuable AI system in science will generate hypotheses, molecules, or experiments.
Some will connect the decision workflow around them.
Perceptic[10] is best understood this way. Not "AI that discovers the drug," but AI that connects the messy reality around drug development: asset scouting, indication selection, clinical data analysis, scientific evaluation, decision context, and organizational memory.
This may sound less glamorous than molecule generation. It may be more immediately useful.
Pharma R&D is not just bottlenecked by scientific imagination. It is bottlenecked by fragmentation:
- data lives in different systems;
- teams work in silos;
- evidence is scattered across papers, dashboards, trial records, internal reports, and expert judgment;
- major decisions are often made by stitching together incomplete context.
This is not a scientific discovery loop in the narrow sense. It is an organizational decision loop.
That distinction matters.
In pharma, the feedback does not close only against biology. It also closes against portfolio strategy, clinical operations, regulatory constraints, partner diligence, and market timing.
A pharma operating system is a bet that the next productivity gain comes from connecting the workflow, not only improving a model.
Cross-cutting constraint: evaluation
Evaluation is not a layer.
It cuts across every loop.
If AI systems generate more hypotheses, we need to know which ones are sound. If they write more papers, we need to know which claims are supported. If they design more proteins, we need to know whether the design works outside the benchmark. If they run more computational experiments, we need to know whether the setup was meaningful. If they summarize more evidence, we need to know what was missed, distorted, or overclaimed.
This is why SoundnessBench[15] matters.
It asks a simple but critical question: can AI judge whether a research proposal is scientifically sound?
The answer is not yet comforting. Current models can look convincing while missing methodological weaknesses. They can reward plausible ideas. They can display optimism bias. They can scale the appearance of rigor without necessarily scaling rigor itself.
RefusalBench[14] matters from another angle.
As AI systems become orchestration layers for biology, they need to know when to help, when to refuse, and when a legitimate research request is being blocked by a crude safety policy. The failure modes are symmetric:
- too permissive, and we scale dual-use risk;
- too restrictive, and we block legitimate research;
- too shallow, and we optimize for compliance theater rather than scientific judgment.
This is the core risk of AI for science.
Not that it produces nonsense. That would be easy to reject.
The risk is that it produces work that looks scientific enough to pass quickly through overloaded human systems.[16]
The emerging map
The sector is not converging around one "AI scientist."
It is decomposing the research loop into different kinds of systems, defined by the cost and fidelity of their feedback:
- code-closed systems for computational experimentation;
- literature-closed systems for hypothesis generation;
- simulation-closed systems for cheap but potentially low-fidelity prioritization;
- protein- and molecule-closed systems for biological design;
- patient-, regulator-, and market-closed systems for therapeutic reality;
- workflow-closed systems for pharma decision intelligence;
- cross-cutting evaluation systems for trust, safety, and scientific rigor.
This is not one market. It is the decomposition of the scientific process into software-addressable loops.
The hard question is no longer whether AI can generate scientific work. It is where the loop closes, and how much we should trust the feedback: against code, literature, simulations, proteins, cells, animals, patients, regulators, and markets.
Cost and fidelity do not move together.
That is the whole point.
Some loops are cheap and reliable. Some are cheap and misleading. Some are expensive but decisive. Some are expensive and still noisy.
The winners will not be the systems that generate the most ideas. They will be the systems that improve the rate at which good ideas become validated knowledge.
That is the real promise of AI for science.
Not infinite generation.
Better scientific judgment at scale.
References
- Scannell, J. W. et al. "Diagnosing the decline in pharmaceutical R&D efficiency." Nature Reviews Drug Discovery, March 2012.
- Gottweis, J. et al. "Accelerating scientific discovery with Co-Scientist." Nature, May 2026.
- Google DeepMind. "Co-Scientist: A multi-agent AI partner to accelerate research." May 2026.
- Lu, C. et al. "Towards end-to-end automation of AI research." Nature, March 2026.
- Sakana AI. "The AI Scientist: Towards Fully Automated AI Research." 2026.
- Biohub. "Biohub releases a world model of protein biology." May 2026.
- Isomorphic Labs. "Isomorphic Labs announces $2.1B Series B investment round." May 2026.
- Incyte and Genesis Molecular AI. "Incyte and Genesis expand molecular AI collaboration to accelerate drug discovery." May 2026.
- Index Ventures. "Inherent: Designing for Discovery." May 2026.
- Air Street Capital. "Introducing Perceptic." May 2026.
- Y Combinator. "CellType: The agentic drug company." February 2026.
- Y Combinator. "BioStack Platforms." May 2026.
- Insilico Medicine and Human Longevity. "Collaboration to co-develop a foundation model for longevity science." May 2026.
- Weidener, L. et al. "RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts." arXiv, May 2026.
- Ho, S.-T. et al. "SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?" arXiv, May 2026.
- Nature. "Why AI cannot do good science without humans." Editorial, Nature, May 2026.