Zendoric
← Back to the day · July 22, 2026

James Zou (Stanford) runs a 'virtual lab' where AI agents generate hypotheses and sign biomedical papers

🕒 Published on Zendoric: July 22, 2026 · 01:59

James Zou, a professor of biomedical data science at Stanford, has built a virtual biotech company and a virtual lab where AI agents generate hypotheses, design experiments and even draft papers with minimal human oversight. The twist isn't that AI answers questions: it's that it's beginning to behave like a researcher.

By Stanford Medicine · July 21, 2026.

James Zou, associate professor of biomedical data science at Stanford, is no longer just the principal investigator of his group: he is, in his own words, in charge of a "virtual biotech company" and a "virtual lab" run by artificial intelligence agents. As he explains in an interview published by Stanford Medicine, his team has developed several agentic AI systems —models capable of reasoning, setting goals and acting autonomously, not just conversing— that take part in almost every phase of research: from shaping a project's idea to reviewing the final manuscript before submitting it to a journal.

The technical distinction Zou draws is clear: if a chatbot is the "brain" that thinks and speaks, an agent adds "arms and legs". An agent still relies on an underlying language model, but it can also use external tools —a web browser, protein molecular modeling software— and thereby make decisions and execute tasks on its own, without a human supervising every step. In some projects in Zou's lab, as he describes it, it is the agent that directly conceives the research idea, analyzes the data and writes the paper, with the human limited to a final review.

The catalog of projects he mentions illustrates how far the delegation has gone: the "virtual lab" acts as a research-director agent that formulates and answers scientific questions; the "virtual biotech company" digitally replicates a biotech to explore potential drug targets; and "Paper2Agent" turns every published scientific paper into its own agent, capable of answering questions about that study and of "conversing" with the agents of other papers. Zou has even experimented with giving agents personalities —recreating Einstein or Feynman as AI characters so they debate scientific questions among themselves—, a use more exploratory than operational, but one that points to where the field wants to take this: agents with their own role and judgment, not just executors of instructions.

Zou is also explicit about the limits, and here it is worth attributing the caution exactly as he frames it: autonomous agents "still make mistakes", and it is "crucial" that human experts verify and validate any finding before accepting it as good. His central argument is physical, not just methodological: agents cannot operate in the real world yet, so any hypothesis or drug candidate they generate must, without exception, go through a real lab before being taken seriously. It is a distinction worth underscoring against easy enthusiasm: an agent that "discovers" something in simulation has discovered nothing until someone verifies it in a Petri dish.

This connects with something we have already noted in analyzing the wave of agentic AI in other sectors: the barrier to entry is not so much the agent's autonomy as the governance of what it produces. Here the pattern repeats in its scientific version. Biomedical science has been limited for decades by the human bottleneck —there are more plausible hypotheses than researchers and hours to test them—; if agents truly multiply the number of hypotheses generated and prioritized per hour, the ceiling of what is possible shifts. But the bottleneck does not disappear, it moves: from "generating ideas" to "validating which of those ideas are not plausible hallucinations". The labs that win this decade will not be those that deploy the most agents, but those that best know how to filter their output.

The underlying reading, the one that connects with our long-term thesis, is what matters most here: if biomedical science is one of the areas where agentic AI can pay off most —because lab work combines exactly what these systems do well (reading literature, generating hypotheses, designing experiments) with what a human does better (validating in the physical world)—, then accelerating that cycle is accelerating, literally, the pace at which therapeutic targets and drug candidates are identified. It is not an immediate promise of cures; it is a promise of fewer years lost per failed hypothesis that an agent discards before a human spends months in the lab chasing it. That is exactly the kind of quiet acceleration —not a "miracle cure" headline, but shorter and cheaper discovery cycles— that sustains, years out, the thesis that AI can bring us closer to eradicating diseases that are untreatable today. Zou himself is cautious about timelines and promises none of that; the caution is his, the long-term projection is our reading.

The short-term risk should not be lost sight of either: the more agents take part in the scientific production chain —from the hypothesis to the manuscript review—, the greater the temptation to cut human oversight to gain speed, right when Zou insists that oversight remains indispensable. Biomedical science, unlike a chatbot that gives an incorrect answer with no further consequence, has a history of fraud and replication failures that already cost it dearly before generative AI existed; introducing agents that write papers semi-autonomously without reinforcing verification mechanisms in parallel does not accelerate science, it contaminates it faster. The competitive advantage of the coming years will not be having the most autonomous agent, but the most rigorous validation process around it.

🔗 Related on Zendoric

Sources & references