← Back to list

Stanford Just Turned Scientific Papers Into AI Agents That Talk to Each Other.

The experiment started as a way to make research more accessible. What came out of it was something the researchers didn’t fully…

Toyez in Level Up Coding · 2026-10-02 14:48 · 61 claps · 5.6 min read paywalled
#artificial-intelligence #machine-learning #research #data-science #llm
Open on Medium ↗
Wiki topics: LLM · Large Language Models AGT · AI Agents ML · Machine Learning AI · AI · General EDU · Education & Learning 🔬 · Science · General

Stanford Just Turned Scientific Papers Into AI Agents That Talk to Each Other. One Found Something Nobody Had Reported Before.

The experiment started as a way to make research more accessible. What came out of it was something the researchers didn’t fully anticipate.

Photo by Scott Graham on Unsplash

Photo by Scott Graham on Unsplash

Since 1665, a scientific paper has been the same thing. A static record. Words on a page, written by people for other people to read. You cite it. You summarize it. You read it slowly and try to extract the method from the prose. The paper itself does nothing.

That is what James Zou’s lab at Stanford Medicine just changed. And the way they changed it is worth understanding precisely, because the most interesting result from the experiment was not what they set out to find.

What Paper2Agent actually does

The system is called Paper2Agent. It takes a scientific manuscript, including the text, figures, data, and associated code, and converts it into an interactive AI agent. The agent is not a chatbot that summarizes the paper. It reproduces it.

A team of worker agents reads the manuscript and then attempts to run the research from scratch inside a virtual environment. They simulate the experiments. They capture the methods, the reagents, the setup, the execution. Everything a reader would have to manually reconstruct from the methods section gets encoded into the agent’s knowledge during that reproduction process.

The resulting agent can answer questions about the work. It can apply the paper’s methods to new datasets. And it can do something a PDF cannot do at all.

It can talk to other paper agents.

The infrastructure underneath this uses MCP, or Model Context Protocol, which has become a standard way for AI agents to share structured knowledge. Each section of the paper gets stored in a separate file inside the agent’s knowledge base, organized the way a filing system works. Other agents can query it in natural language and get back responses that reflect the full context of the original research, not just what appears in the abstract.

The discovery that came from a conversation between two unrelated papers

Zou’s team demonstrated what agent-to-agent collaboration actually produces by taking two papers that had nothing obvious to do with each other.

One described a computational tool for predicting how genetic mutations affect the genome. The other was a genome-wide association study on the risk of developing ADHD. Different research groups. Different questions. Different datasets. The kind of papers that would never end up in the same literature review because nobody would think to look for a connection between them.

When the two paper agents were given access to each other and allowed to collaborate, the mutation prediction agent applied its methodology to the ADHD dataset. It flagged a molecular variant near a gene called MPHOSPH9 that appeared to be associated with increased ADHD risk.

Zou said that connection had not been previously reported.

That is not retrieval. The information did not exist in either paper. It emerged from the conversation between two agents that represented different bodies of knowledge and were able to apply one to the other in a way no human had thought to do.

Why this is harder than it sounds

The obvious objection is that this is just a very sophisticated literature synthesis tool. Connect any two datasets and you can find correlations. What makes this different from running a query across a database?

The difference is in what the agents actually encode. A database query finds what is in the data. A Paper2Agent instance encodes how the research was done, not just what it found. The worker agents that build the paper agent are not extracting text. They are reproducing the experiment. They capture the judgment calls, the experimental setup choices, the methodology in its executable form.

That is why the collaboration produced something novel. The mutation prediction agent did not just find a correlation in the ADHD dataset. It applied a methodology developed for a completely different purpose to a completely different dataset and found a result that the original authors of neither paper had looked for.

Zou is careful about what this means. The MPHOSPH9 finding is a candidate for further investigation, not a confirmed discovery. The agent-to-agent collaboration surfaces connections worth examining. It does not validate them. Everything that comes out still needs to be tested experimentally before it becomes science.

But the finding demonstrates something real about the potential scale of this approach. Zou’s team selected these two papers and paired them deliberately. The eventual goal is something he describes as manuscript speed dating at scale: millions of paper agents finding each other, surfacing common ground without human direction, and flagging connections for researchers to investigate.

What the human role becomes

One thing Zou is specific about is that the system is not designed to replace human researchers or obscure whose work produced which finding.

Paper agents represent and extend the original authors’ work. Attribution stays with the humans who did the research. When an agent-to-agent collaboration produces a new finding, the credit traces back to the papers that made it possible, not to the agents that found the connection.

There is also something the system cannot capture automatically. Failed experiments. Judgment calls that never made it into the methods section. The context that lives in a researcher’s head rather than in the manuscript. Zou and the human authors still supply that context in conversational exchanges with the paper agent during the setup process. The agent can ask questions. The authors can fill in what the PDF left out.

This distinction matters practically. Two papers from research groups that have never met might contain methods and data that could inform each other. Before Paper2Agent, finding that connection required one of the researchers to read the other paper, recognize the relevance, and reach out. That chain of events is slow, depends on serendipity, and fails most of the time because researchers cannot read everything.

With paper agents, that overlap can surface without anyone having to notice it first. The connection gets flagged. Researchers can then decide whether it is worth pursuing.

What 91.2% accuracy on 100 papers actually means

The benchmark the team ran to validate Paper2Agent’s reliability involved 100 computational biology papers. For each paper, they built an agent and then asked it benchmark questions drawn from the paper’s content. The system averaged 91.2% accuracy across those questions.

That number is worth contextualizing. It means the agent correctly answered about nine out of ten questions about a paper’s methods, findings, and data. It does not mean the agent is right nine out of ten times about novel claims generated through agent-to-agent collaboration. Those claims require a different kind of validation: experimental testing of what the collaboration surfaces.

The accuracy benchmark establishes that the individual paper agents are reliable representations of their source material. The downstream discoveries those agents might find together are a separate question, and one that the MPHOSPH9 finding is the beginning of an answer to, not the end.

The practical shape of what comes next

More than 100 paper agents exist right now inside Zou’s lab. The team is actively working on how to scale that to thousands and eventually millions. Millions of papers are published every year. The eventual infrastructure would need to handle agents finding each other, initiating conversations, and flagging results without a human directing each pairing.

There are real challenges in that. At scale, agents will find many candidate connections. Most will not survive experimental validation. Building the filtering and prioritization layer that helps researchers focus on the connections worth investigating, rather than drowning them in spurious correlations, is where a lot of the remaining engineering work lies.

There are also questions about how agent-to-agent collaboration gets governed at scale. What constraints should exist on how agents apply one paper’s methods to another paper’s data? What disclosure is owed to the original authors when their work is extended in ways they did not initiate? Zou notes that the parameters under which agents collaborate should be closely guided and monitored for safety and ethical research.

None of those questions have clean answers yet. What the September 16 Nature paper establishes is that the basic capability is real: a paper can become an active embodiment of its own knowledge rather than a static record of it, and two such embodiments can find something that neither paper contained.

That is a meaningful change in what scientific publishing can be. The question now is how to build the infrastructure that makes it work at the scale where it becomes genuinely useful.


메타데이터
post_id
d6ad43e0bb13
slug
stanford-just-turned-scientific-papers-into-ai-agents-that-talk-to-each-other-d6ad43e0bb13
url
https://levelup.gitconnected.com/stanford-just-turned-scientific-papers-into-ai-agents-that-talk-to-each-other-d6ad43e0bb13
canonical_url
https://levelup.gitconnected.com/stanford-just-turned-scientific-papers-into-ai-agents-that-talk-to-each-other-d6ad43e0bb13
author_url
https://medium.com/@toyezyadav
status
ok
fetched_at
2026-10-03 09:38:59