· via dev.to (home feed)
Stanford's Paper2Agent turns 100 research papers into runnable MCP agents
Stanford researchers converted 74 of 100 computational biology papers into validated MCP agents that beat a Claude-plus-repo baseline and surfaced findings no single paper contained.

From papers to agents
A Stanford team has published an automated pipeline that converts research papers into agents researchers can query in plain English. According to a write-up on dev.to, the system, called Paper2Agent, was developed by Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard and James Zou, and appeared in Nature on September 16, 2026. Run across 100 computational biology papers, it produced 74 working agents and extracted 599 tools, of which 593 passed automated validation.
The motivation is straightforward: a paper describes a method but does not run it. Anyone wanting to apply published code to their own data must locate the repository, resolve dependency conflicts, guess which notebook implements the method and manually map figures back to numbers.
How the pipeline works
Paper2Agent handles the setup end to end in six stages: it finds the repository, notebooks and documentation linked from a paper; builds an isolated environment with pinned dependencies; collects the tutorials that demonstrate the methods; runs them and stores the outputs as ground truth; wraps the runnable components as parameterized tools with JSON schemas; and packages everything as a Model Context Protocol (MCP) server that any chat agent can call.
A validation gate sits between tool extraction and assembly. As reported by dev.to, numeric outputs must match the paper's own results within 3%, generated figures are compared using perceptual hashing, and tools are locked so the calling model cannot improvise new behaviour. Adversarial testing reportedly produced a 100% correct rejection rate on out-of-scope queries, and any tool that fails to reproduce the paper's outputs is discarded.
The benchmark numbers
On 300 benchmark questions derived from tutorials and spread across all 74 agents, the agents scored 91.2% (plus or minus 1.6) against 82.7% (plus or minus 3.4) for the baseline of giving Claude the same repositories as raw files. The same model family was used in both setups; the dev.to post attributes the gap to packaging, arguing that callable, verified tools outperform direct code access.
Costs are modest. The AlphaGenome agent was built in roughly 45 minutes for about $14 of compute, yielding 22 tools; it scored 98.7% on 15 tutorial queries and 100% on 15 novel questions written by the authors. Beyond biology, 42 execution tasks from 10 other computational papers reached 98.1%, with median runtime 1.9 times faster than the Claude-plus-repo setup and 3.1 times faster than Biomni. Per query, the agents cost $0.20 and 1.6 minutes versus $0.38 and 4.3 minutes for the baseline.
Agents that combine findings
The pipeline also enables something individual papers cannot do. The team built agents from three unrelated papers, an AlphaGenome variant-effect predictor, an MPRA-coupled scCRISPRi screening agent and a CD4+ T cell Perturb-seq agent, and chained their predictions to prioritise GPR137 as a probable causal gene at the psoriasis-associated locus rs887314. A separate Stanford Medicine demonstration paired a mutation-prediction agent with an ADHD genome-wide association study to surface a previously unreported variant near MPHOSPH9 associated with increased ADHD risk. James Zou reportedly envisages millions of such agents finding overlapping work at scale, with discoveries still credited to the original papers and their human authors.
Stated limitations
Twenty-six papers could not be converted, blocked by missing code, unavailable data or environments that would not build, a failure rate the authors reportedly treat as a reproducibility audit rather than a defect. Open-ended hypothesis selection and evidence evaluation remain human tasks. The MCP servers need ongoing maintenance as upstream dependencies drift, making each agent a living artifact rather than a filed-away PDF. And the 100% result on novel queries rests on only 15 questions written by the authors themselves, so independent benchmarks will be the real test. Outside researchers quoted in the coverage, including Olivier Elemento of Weill Cornell and Dongping Chen of the University of Maryland, were positive but stopped short of calling this a replacement for peer review.
Why it matters
Paper2Agent reframes the research paper as a deployable interface rather than a static document. If validated tools can be generated cheaply, around $14 per paper in this run, the marginal cost of actually using published methods collapses, and reproducibility becomes a measurable property: papers that survive the pipeline are demonstrably complete, while those that fail are flagged. The multi-agent results point further, toward networks of paper-derived agents cross-referencing work their original authors would never have connected. The open question is durability, whether the community maintains these artifacts or they decay like the unmaintained repositories they came from.
- #ai-agents
- #mcp
- #reproducibility
- #research
- #llm