Paper2Agent Turns Research Papers Into AI Agents That Can Actually Run the Research
Paper2Agent Turns Research Papers Into AI Agents That Can Actually Run the Research

Paper2Agent Turns Research Papers Into AI Agents That Can Actually Run the Research

Share story

Advertisement

A scientific paper can explain an experiment in exquisite detail and still leave another researcher wondering one frustrating question:

How do I actually run this?

The manuscript may describe the method.

The supplementary files may contain parameters.

A GitHub repository may hold the code.

A tutorial may explain part of the workflow.

Dependencies may require specific versions.

Data formats may be undocumented.

And somewhere between reading the paper and reproducing the result, hours—or days—can disappear.

A new framework called Paper2Agent is trying to change that.

Published in Nature on September 16, 2026, Paper2Agent converts research papers and their associated software into interactive AI agents that can answer technical questions, invoke the paper's computational methods, analyze compatible new data, and expose the original workflow through natural-language interaction.

The idea is much more ambitious than asking a chatbot to summarize a PDF.

Instead of merely explaining what a paper says, a Paper2Agent-derived system can, when the underlying research code supports it, execute the methods described by the paper.

In the researchers' demonstrations, agents built from tools including AlphaGenome, Scanpy, and TISSUE could perform real computational workflows and reproduce results that human researchers obtained by running the original software.

That creates a provocative possibility:

The research paper of the future may not just be something scientists read.

It may be something they can talk to, test, and run.

What If a Research Paper Could Answer Questions?

Scientific publishing has traditionally been static.

A paper contains:

an abstract,

methods,

results,

figures,

discussion,

references,

and perhaps links to supplementary data or software.

The reader has to reconstruct the rest.

If the study introduces a new computational method, using it may require learning an unfamiliar programming interface, installing libraries, identifying compatible software versions, finding example datasets, and translating the written methods into a working sequence of commands.

That is manageable for experts who already know the field.

It becomes a major barrier for everyone else.

Paper2Agent tries to insert an interactive layer between the research and the user.

Instead of asking:

“How did the authors run this analysis?”

a researcher might ask:

“Run this method on my compatible dataset and show me the result.”

The agent can then invoke structured tools derived from the source repository rather than inventing an analysis from scratch.

Nature's accompanying research briefing described the concept as transforming papers into active systems capable of answering questions, applying methods to new data, and collaborating with other paper-derived agents.

That is a significant shift in what a scientific publication could become.

What Is Paper2Agent?

Paper2Agent is an automated multi-agent framework for converting scientific research packages into interactive AI agents.

The system does not simply read the prose of the paper.

It examines the surrounding research ecosystem:

the manuscript,

supplementary material,

code repository,

tutorials,

dependencies,

data resources,

and executable workflows.

Multiple specialized agents then analyze those materials and identify useful scientific functions that can be exposed as tools.

The resulting system is connected through the Model Context Protocol, or MCP, allowing an AI assistant to call the paper's software through standardized interfaces.

In simpler terms:

Paper2Agent tries to turn a research repository into a toolset an AI can use correctly.

It Is More Than a Paper-Summary Chatbot

This distinction is crucial.

Imagine uploading a paper about a genomic algorithm to a general chatbot.

You ask:

“Can you analyze this genetic variant using the method in the paper?”

The chatbot might:

summarize the algorithm,

write approximate Python code,

guess which parameters should be used,

or hallucinate an output.

Even if the explanation sounds convincing, the answer may never have passed through the actual scientific software.

Paper2Agent takes a different approach.

When an executable agent can be constructed, the language model is not supposed to replace the scientific method.

It acts as an interface to it.

The underlying software still performs the substantive calculation.

The AI helps decide:

which validated tool should run,

what parameters it needs,

in what order tools should be called,

and how the resulting output should be explained.

That is a much stronger foundation than asking a language model to reproduce a method from memory.

Why the Associated Code Matters So Much

Paper2Agent works best when the paper is accompanied by usable software.

A strong research repository may include:

well-documented functions,

tutorial notebooks,

example datasets,

dependency files,

tests,

clear installation instructions,

model checkpoints,

and explicit version information.

That gives Paper2Agent material it can inspect, execute, test, and wrap into tools.

A paper with no working code creates a very different situation.

The framework can still expose manuscript content and structured resources in some cases, but it cannot magically turn a non-existent implementation into a perfectly validated scientific method.

This limitation showed up clearly in the large-scale evaluation.

Of 100 computational-biology papers processed automatically, 74 were successfully converted into executable agents. Failures were associated with problems including missing code, unavailable data or models, broken dependencies, and scripts that were not sufficiently generalizable.

That result is important because it reveals what Paper2Agent really depends on:

good computational science remains the foundation.

AI cannot repair every broken research package.

How Paper2Agent Turns a Study Into an AI Agent

The framework can be understood as a pipeline.

Step 1: Inspect the Research Package

Paper2Agent begins by examining the relevant materials.

The system tries to determine:

What does this research package actually do?

What functions are useful?

What inputs are supported?

What outputs are produced?

What software dependencies are required?

Which tutorials demonstrate intended use?

Which parts of the code are reusable?

This stage matters because repositories are often written for researchers rather than machines.

A human might infer that three scripts must be executed in a particular sequence.

An automated system needs that relationship made explicit.

Paper2Agent attempts to reconstruct those relationships.

Step 2: Convert Scientific Functions Into Tools

Once useful capabilities are identified, the framework wraps them into structured callable tools.

A tool might correspond to something like:

running quality control,

scoring a genetic variant,

clustering single-cell data,

calculating an uncertainty interval,

or loading a particular dataset.

Those functions are exposed through MCP.

What Is MCP?

The Model Context Protocol is a standardized way for AI systems to interact with external tools and data sources.

An easy analogy is a universal adapter.

Instead of every AI assistant requiring a completely custom integration for every scientific program, MCP defines a common structure through which tools can describe:

what they do,

what inputs they expect,

and what they return.

Paper2Agent uses this approach to make research software callable by an AI agent.

That separation is important.

The AI handles natural-language interaction.

The scientific tool handles computation.

Step 3: Generate Tests and Validate the Tools

This may be the most important part of the framework.

Paper2Agent does not simply generate wrappers and assume they work.

It runs tests.

If a generated tool fails, the system can inspect the problem and refine the implementation.

The framework checks whether the tool can:

accept expected inputs,

execute correctly,

return usable outputs,

and reproduce intended behavior.

In the 100-paper computational-biology evaluation, Paper2Agent proposed 599 tools, and 593 passed automated validation.

That validation layer distinguishes the framework from casual AI-generated code.

A language model can produce plausible software very quickly.

Scientific infrastructure requires something stronger:

software that actually runs.

Step 4: Connect the Tools to a Conversational Agent

Once the MCP tools exist, an MCP-compatible AI client can use them.

The user can then ask questions in ordinary language.

For example:

“Preprocess and cluster this single-cell dataset.”

The agent determines which validated tools are relevant and invokes the underlying workflow.

In the Scanpy demonstration, users could provide a dataset path and the derived agent would execute a preprocessing and clustering pipeline including quality control, normalization, feature selection, dimensionality reduction, graph construction, clustering, and cell-type annotation.

The interface feels conversational.

The computation underneath remains structured.

The Scientific Demonstrations

The Paper2Agent team tested the framework using several well-known computational-biology tools.

Three examples make the concept especially easy to understand.

AlphaGenome: Asking Genomic Questions Conversationally

AlphaGenome is a genomic prediction system designed to model how DNA sequence changes may affect molecular biology.

Paper2Agent converted its software capabilities into an agent that could interpret genetic variants through conversational queries.

Instead of manually constructing a script to call the model, identify a genomic coordinate, request particular molecular predictions, extract the correct output, and interpret it, the researcher could ask the agent to perform the workflow.

In controlled benchmarks, the AlphaGenome-derived Paper2Agent system performed very strongly on both tutorial-derived and novel queries.

But an important caveat remains:

This does not turn the agent into a clinical geneticist.

A computational prediction about a variant is not equivalent to a medical diagnosis.

Biological interpretation still requires expert judgment and, frequently, independent experimental validation.

Scanpy: Single-Cell Analysis Without Memorizing the Entire API

Scanpy is widely used for single-cell gene-expression analysis.

A typical workflow may involve:

loading data,

quality control,

filtering,

normalization,

feature selection,

dimensionality reduction,

neighbor-graph construction,

clustering,

and marker-gene analysis.

Researchers familiar with Scanpy can perform these operations directly.

A newcomer can spend considerable time learning the correct order and parameters.

Paper2Agent created seven validated tools for a focused Scanpy preprocessing and clustering workflow. The authors reported that the conversion took about 45 minutes and roughly US$13 using their setup.

The derived agent was then tested on multiple datasets.

It produced results comparable with human researchers following official Scanpy tutorials and adapted some parameters according to dataset characteristics.

That is precisely the type of task where an interactive paper agent could be useful.

The expert software remains in control of the analysis.

The agent lowers the interface barrier.

TISSUE: Spatial Transcriptomics With Uncertainty

TISSUE is a method related to spatial transcriptomics.

Spatial transcriptomics attempts to measure gene activity while preserving information about where those signals occur in tissue.

That spatial information matters because biology depends heavily on location.

A cell beside a tumor boundary may behave differently from a similar cell elsewhere.

TISSUE also focuses on uncertainty-aware analysis.

Paper2Agent generated a TISSUE MCP agent and tested it against analyses performed by human researchers using the same data and tutorial workflow. The researchers reported matching results in their reproducibility evaluation.

This is a useful example because scientific AI should not merely produce an answer.

It should also preserve uncertainty when the underlying method provides it.

How Well Did Paper2Agent Perform?

The largest evaluation is one of the most interesting parts of the research.

The team processed three collections:

100 computational-biology papers

26 data- and discovery-focused papers

and

10 computational papers outside biology.

The 100-paper biology benchmark provides the most quoted result.

About 91% Accuracy Across 300 Benchmark Questions

Among the 74 computational-biology papers that could be successfully agentified, Paper2Agent was tested on 300 tutorial-derived benchmark questions.

Using Sonnet 4, the Paper2Agent system achieved an average accuracy of approximately 91.2% ± 1.6%.

The paper reported higher performance than Claude Code given direct access to the repositories under the tested setups.

That is an impressive number.

It is not 100%.

And the distinction matters.

A 91% benchmark score means roughly nine out of ten benchmark items were handled correctly under that evaluation.

It does not mean:

91% of scientific conclusions will be correct,

91% of all papers can be converted,

or the system is 91% reliable in every real-world research context.

Benchmarks measure particular tasks under particular conditions.

Seventy-Four of 100 Papers Could Be Agentified

This number may be just as important as the accuracy figure.

Paper2Agent attempted to process 100 computational-biology papers automatically.

Only 74 yielded executable agents.

That means the success rate of the conversion process itself was roughly three quarters in that sample.

The failures reveal a fundamental truth about executable science:

If the original research package is incomplete, brittle, undocumented, or dependent on missing artifacts, an AI layer cannot guarantee reproducibility.

Paper2Agent may expose the problem more quickly.

It cannot always solve it.

Results Beyond Computational Biology

The researchers also tested Paper2Agent on 10 non-biology computational papers spanning areas including AI, statistics, econometrics, game theory, and other computational disciplines.

Across 42 execution-based tasks, the reported accuracy was about 98.1% over repeated runs.

That result suggests the concept may generalize beyond computational biology.

But ten papers are not enough to claim universal applicability.

Computational papers with well-defined software interfaces are especially favorable candidates for agentification.

A qualitative sociology paper, a clinical trial, or an archaeological field report presents a very different problem.

What Happens When There Is No Executable Tool?

Paper2Agent also includes a resource layer.

This allows papers that cannot become executable agents to remain queryable through structured manuscript, supplementary, and data resources.

In a set of 26 data- and discovery-focused papers, the resource layer achieved strong performance on synthesis questions and outperformed the comparison browser-use system in the study.

This creates two levels of “paper agent.”

One may actually execute validated research code.

Another may primarily expose research materials in a structured way.

Users need to know which they are interacting with.

Otherwise “the paper can answer questions” could sound far more capable than the underlying implementation really is.

Reproducing a Workflow Is Not the Same as Making a Discovery

This distinction may be the most important in evaluating Paper2Agent.

Suppose an agent reproduces a tutorial result.

That can be checked.

The original output is known.

The expected parameters may be known.

The underlying software already exists.

Now imagine asking:

“Analyze this completely new dataset and tell me what biological mechanism explains the result.”

That is a fundamentally harder problem.

The software may execute correctly.

The interpretation can still be wrong.

Scientific discovery requires more than tool execution.

It involves:

experimental design,

domain knowledge,

causal reasoning,

statistical judgment,

recognition of confounders,

understanding of biological plausibility,

and independent validation.

A paper agent can assist.

It does not eliminate those requirements.

Paper Agents Can Even Disagree With the Original Paper

One of the most interesting demonstrations involved the AlphaGenome agent.

When examining one variant, the derived system prioritized SORT1 as the most likely causal gene, while the original publication had emphasized other nearby genes.

The researchers then examined additional evidence and noted that several genes remained plausible, illustrating how difficult causal-gene assignment can be at complex loci.

This is fascinating because it shows that an agentified paper is not necessarily a frozen reproduction of the authors' interpretation.

Once its tools become interactive, researchers can ask new questions.

That is powerful.

It is also risky.

A newly generated computational hypothesis remains a hypothesis.

An AI system producing a different answer from the original paper does not mean it has discovered that the authors were wrong.

Can Research Papers Collaborate With One Another?

Paper2Agent points toward an even more ambitious possibility:

paper-agent ecosystems.

Imagine one agent derived from a genome-association study.

Another exposes a genomic prediction model.

A third queries gene-expression resources.

A fourth performs statistical analysis.

Instead of manually combining each method, an AI system could orchestrate several specialized paper-derived agents.

The Nature paper includes demonstrations of agent collaboration for genomic analysis and hypothesis generation.

In principle, that could transform literature review from:

read → understand → install → rewrite → test

into something closer to:

query → invoke → compare → validate.

But this remains an emerging research direction.

It is not yet a mature scientific operating system.

The Danger of Compounding Errors

If one agent makes a mistake, another agent may consume the incorrect output as though it were valid.

Then a third agent builds on that.

The final answer may look sophisticated precisely because several tools participated.

Yet the error occurred at the beginning.

Multi-agent scientific systems therefore need strong provenance.

Every output should ideally record:

which paper supplied the method,

which version of the code ran,

which data were used,

which parameters were chosen,

which agent invoked the tool,

what warnings occurred,

and whether validation checks passed.

Without that audit trail, agent collaboration could make scientific errors harder—not easier—to diagnose.

Could Paper2Agent Help Science's Reproducibility Problem?

Potentially, yes.

But only part of it.

Reproducibility problems often arise because computational methods are difficult to reconstruct.

A published analysis may depend on:

a specific library version,

an undocumented preprocessing step,

a missing reference file,

a particular operating-system package,

a script executed manually,

or a model checkpoint no longer available.

Paper2Agent can help identify and standardize some of those dependencies.

It can create callable interfaces.

It can execute tutorials.

It can validate generated tools.

It can make workflows easier for outsiders to rerun.

That is meaningful progress.

What Automation Cannot Repair

Paper2Agent cannot fix a scientifically flawed source study merely by making it executable.

If the original data are biased, the agent inherits that problem.

If the statistical method is inappropriate, automated execution makes the mistake faster.

If the repository does not reproduce the paper's results, wrapping it in MCP does not transform it into correct science.

If important data are proprietary or missing, the agent cannot reconstruct them reliably.

If the authors selectively reported favorable results, agentification does not undo publication bias.

This is a crucial principle:

reproducibility and validity are not the same thing.

A flawed method can be perfectly reproducible.

Can the Generated Agent Be Trusted?

Trust should be layered.

You can ask several separate questions.

Did the tool execute successfully?

Did it reproduce the original tutorial?

Was the correct version of the software used?

Did it apply the method to compatible data?

Was the scientific interpretation reasonable?

These are different validation problems.

Automated testing is particularly good at the first few.

The last one often requires expert review.

An executable agent can therefore be more trustworthy than improvised AI-generated code without becoming infallible.

Does the Agent Preserve the Authors' Intent?

Not necessarily.

Paper2Agent infers useful workflows from the source material.

But scientific methods often contain assumptions that are obvious to the authors and poorly documented in the repository.

An automated system might expose a function beyond the conditions under which the authors intended it to be used.

That suggests a future need for two categories:

author-approved paper agents

and

community-generated paper agents.

An official agent reviewed by the paper's authors could carry a different level of trust from one automatically generated by a third party.

Versioning Could Become a Major Problem

Traditional papers are static.

Software is not.

A paper published in 2026 may depend on:

Python 3.x,

a particular machine-learning library,

a specific API,

a model version,

and an external data resource.

Five years later, several may have changed.

The paper still exists.

The agent may no longer work.

This raises an unexpectedly difficult question:

Which version of the paper agent represents the paper?

The version matching publication day?

The most recent bug-fixed version?

An author-maintained version?

A community fork compatible with new dependencies?

A fully containerized archival version?

Scientific publishing has not had to answer this question at scale because PDFs do not break when Python changes.

Executable publications can.

Retractions Create an Even Bigger Problem

Suppose a paper is retracted.

The PDF receives a retraction notice.

What happens to its agent?

If copies of the derived MCP server are running elsewhere, they may continue executing the method indefinitely.

Future infrastructure will need a way to propagate:

corrections,

expressions of concern,

retractions,

security warnings,

and version updates

to every derived agent.

Otherwise an obsolete agent could outlive the scientific credibility of its source.

Open Science Does Not Mean Everything Is Free to Reuse

Paper2Agent's code is publicly available, and the Nature paper provides links to the project implementation and demonstration MCP servers.

But converting scientific work into agents raises complicated licensing questions.

A research package may contain several legal layers:

the article,

the code,

the dataset,

trained models,

figures,

third-party dependencies,

and external APIs.

Each may have a different license.

An open-access article does not automatically make all associated software and data unrestricted.

Anyone building or distributing paper agents at scale will need to preserve those distinctions.

Attribution Will Matter More, Not Less

If several paper agents collaborate, who receives credit?

Suppose:

Paper A provides the dataset.

Paper B provides the statistical method.

Paper C provides a biological model.

Paper D supplies a visualization library.

The final AI-generated analysis should not erase those contributions.

A mature paper-agent ecosystem will need machine-readable citation and provenance.

Every output should ideally preserve attribution to:

paper authors,

software maintainers,

dataset creators,

model developers,

and other dependencies.

Otherwise interactive research could make scientific reuse easier while making scientific credit harder.

Could Publishers Commercialize Interactive Papers?

It is easy to imagine a future journal page containing:

Read Paper

Download Data

View Code

Ask Paper

Run Method

Interactive agents could become a publishing product.

That could be useful.

It could also create a new kind of paywall.

Today a researcher may be blocked from reading an article.

Tomorrow they might read the article but have to pay to execute its agent.

The alternative is community-operated infrastructure in which executable research remains open and portable.

Which model wins will depend not only on technology but on publishing economics.

Who Could Benefit Most From Paper Agents?

Working Scientists

The immediate use case is straightforward.

A scientist encounters an unfamiliar method.

Instead of spending hours learning the entire codebase, they could use a validated paper agent to test whether the method fits their data.

That lowers the cost of exploring new techniques.

It could also make cross-disciplinary research easier.

A neuroscientist may want to use a specialized statistical tool without becoming an expert in its software interface.

The agent can reduce that friction.

Students and Educators

Paper agents could transform scientific education.

Instead of merely reading:

“Changing parameter X affects clustering resolution,”

a student could change the parameter and observe what happens.

Interactive methods could make scientific papers behave more like laboratories.

But education also creates a risk.

If students can ask the paper agent for answers without understanding the method, convenience may replace learning.

Good educational use would therefore emphasize experimentation and interpretation rather than one-click output generation.

Journal Reviewers

Peer review could become more computationally rigorous.

A reviewer might be able to:

run the author's pipeline,

change parameters,

test a sample dataset,

verify dependencies,

and identify broken workflows

before publication.

Paper2Agent itself cannot guarantee research integrity.

But systems like it could make computational verification cheaper.

That would be valuable.

Research-Integrity Teams

Journals and institutions increasingly confront papers whose code cannot be reproduced.

A standardized agentification process could reveal problems early.

If a repository cannot be converted because critical files are missing, that failure itself is informative.

In other words:

sometimes the most valuable output of Paper2Agent may be discovering that a paper cannot become an agent.

Non-Experts

This is the most exciting and potentially dangerous audience.

Paper agents could give non-programmers access to advanced scientific methods previously available only to specialists.

That democratization could be transformative.

It could also produce confident misuse.

A person without genetics training may treat a model prediction as a medical conclusion.

A business user may run an economic model outside its intended assumptions.

A student may interpret a cluster as a biological discovery.

Ease of execution does not produce expertise.

Interfaces therefore need guardrails explaining when expert interpretation is required.

What Paper2Agent Does Not Mean

The excitement around interactive research can easily outrun what the technology actually demonstrates.

Several misconceptions are worth avoiding.

It Does Not Make Traditional Papers Obsolete

Scientists still need a stable scientific argument.

The paper explains:

why the study was conducted,

what evidence supports the claim,

how the authors interpret the results,

what limitations exist,

and how the work relates to previous literature.

An agent may help execute the method.

It cannot replace the need for an archived scholarly record.

The likely future is paper plus agent, not paper versus agent.

It Is Not an Autonomous Scientist

Running validated tools is not equivalent to doing science independently.

A scientist has to decide:

which question matters,

whether the dataset is appropriate,

whether the controls are sufficient,

whether assumptions are violated,

what alternative explanations exist,

and what experiment should happen next.

Agents may increasingly assist with these tasks.

But executing a workflow is only one part of research.

Ninety-One Percent Accuracy Does Not Mean Universal Reliability

The much-discussed 91.2% result came from a particular benchmark across agentifiable computational-biology papers.

Performance will vary according to:

repository quality,

documentation,

discipline,

task complexity,

software design,

and how far the query moves beyond the original use case.

Computational biology is also unusually compatible with this approach because so much modern work is already expressed through code.

Fields dominated by field observations, laboratory manipulation, qualitative interpretation, or proprietary instruments may be harder to agentify.

The Future of Scientific Publishing Could Be Executable

Imagine opening a paper in 2030.

Beside the abstract is a panel showing:

Paper version: 2.1

Code environment: archived

Agent validation: passed

Reproducibility certificate: verified

Last dependency test: 12 days ago

Retraction status: clear

You ask:

“Can this method analyze my dataset?”

The agent checks the schema.

It says yes.

You upload the data.

It runs the published workflow.

Every tool call appears in an audit log.

Every output links back to the relevant method.

Parameters are preserved.

Citations are attached automatically.

Limitations appear alongside the result.

That kind of publishing infrastructure now seems technically conceivable.

Paper2Agent is an early attempt to move in that direction.

But Executable Science Will Need New Standards

If paper agents become common, scientific publishing will need standards that do not currently exist.

Among them:

Agent identity — Which publication and version does this agent represent?

Versioning — What happens when code changes?

Reproducibility — Has the agent independently reproduced known results?

Security — Is executing the underlying code safe?

Provenance — Which tools, data, and parameters produced the answer?

Attribution — Which papers and software projects should be cited?

Correction propagation — What happens if the paper is corrected?

Retraction propagation — Can an invalidated paper agent be clearly flagged?

Archiving — Can the system still run ten years later?

Without standards like these, interactive papers could become convenient but scientifically fragile.

Security May Become an Underappreciated Problem

A paper agent executes software.

That immediately creates risks a PDF does not have.

Research code can contain:

unsafe file operations,

network calls,

untrusted dependencies,

credential requirements,

or vulnerabilities.

A large ecosystem of executable papers would therefore need code isolation and security review.

The scientific question might be innocent.

The software environment executing it still needs protection.

This is one reason containerization, restricted permissions, reproducible environments, and sandboxing could become central parts of future paper-agent infrastructure.

The Decisive Test Is Not Whether the Demo Looks Impressive

The decisive test will be whether independent scientists find these systems genuinely dependable.

Several questions matter more than flashy demonstrations.

Can another laboratory reproduce the published benchmark?

Do agents still function after dependency updates?

Can they detect when incoming data violate assumptions?

Do they expose failed analyses clearly rather than improvising?

Do they preserve the exact parameters used?

Can researchers distinguish validated execution from AI-generated interpretation?

Do they save enough time to justify the infrastructure?

Most importantly:

Do they reduce scientific error—or merely make analysis faster?

That is the standard Paper2Agent ultimately has to meet.

Research That Can Be Used, Not Just Read

Scientific publishing has spent centuries perfecting the document.

Paper2Agent asks what comes after the document.

Its answer is a provocative one:

make the paper executable.

The framework does not eliminate the need to understand research.

It does not guarantee reproducibility.

It does not turn every paper into a reliable agent.

It does not make AI a scientist.

What it does is connect several technologies that previously existed separately:

research papers,

scientific code,

automated software testing,

large language models,

and standardized tool interfaces.

The result is a publication that can potentially expose its methods directly through conversation.

That changes the relationship between reader and paper.

Instead of only asking:

“What did the authors do?”

the reader can increasingly ask:

“Can I run it?”

“Can I test it on my data?”

“Can I change the parameters?”

“Can another paper's method work with this one?”

“Can I reproduce the result?”

Those are powerful questions.

Paper2Agent does not yet guarantee trustworthy answers to all of them.

But it demonstrates that the scientific paper no longer has to remain a passive endpoint.

It can become an interface.

A tool.

Perhaps eventually, a collaborator.

The real breakthrough will not come when every research paper has a chatbot attached.

It will come when researchers can trust an interactive paper because every computation remains transparent, attributable, versioned, reproducible, and independently verifiable.

That is a much harder goal.

Paper2Agent may have shown one possible path toward it.

Frequently Asked Questions

What is Paper2Agent?

Paper2Agent is an automated multi-agent framework that transforms research papers and associated software into interactive AI agents capable of exposing paper resources and, when suitable code exists, executing validated scientific tools.

When was Paper2Agent published?

The peer-reviewed Paper2Agent study was published in Nature on September 16, 2026.

Who created Paper2Agent?

The Nature paper was authored by Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, James Zou, and collaborators.

How does Paper2Agent turn a paper into an AI agent?

It examines the paper and associated codebase, identifies usable scientific functions, converts them into MCP-compatible tools, generates tests, validates those tools, and makes them callable by an AI assistant.

What is an interactive research paper?

It is a publication whose methods or resources can be queried programmatically rather than only read as static text. In Paper2Agent's implementation, the agent can expose information and sometimes execute the paper's computational methods.

What is the Model Context Protocol?

MCP is a standardized protocol that allows AI systems to interact with external tools and data sources. Paper2Agent uses MCP servers to expose scientific software to conversational agents.

Does Paper2Agent simply summarize a PDF?

No. Its main innovation is connecting AI interaction with the paper's associated software and resources. When the code can be agentified, the system can execute actual scientific workflows.

Can Paper2Agent reproduce published results?

In several demonstrated cases, including Scanpy and TISSUE workflows, Paper2Agent-derived agents reproduced results comparable with those obtained by human researchers following the original workflows.

How accurate was Paper2Agent?

Across 300 tutorial-derived questions from successfully agentified computational-biology papers, the study reported approximately 91.2% average accuracy under the tested configuration.

Did Paper2Agent successfully convert all 100 biology papers?

No. 74 of 100 computational-biology papers were successfully agentified into executable tools. Missing code, unavailable artifacts, dependency problems, and non-generalizable scripts were among the major failure modes.

How many tools did it generate?

Across those 100 computational-biology papers, Paper2Agent proposed 599 tools, with 593 passing automated validation.

Does Paper2Agent work outside biology?

The paper tested it on 10 non-biology computational papers spanning several disciplines, with strong results on the study's execution benchmark. That is promising, but it does not yet prove equal performance across all scientific fields.

Does Paper2Agent work without original code?

Its structured resource layer can still make paper content and associated materials queryable, but reliable executable methods are much harder to construct when usable code is missing.

What is the AlphaGenome Paper2Agent example?

The researchers built an agent that could invoke AlphaGenome tools to analyze genomic variants through natural-language requests.

What is the Scanpy Paper2Agent example?

Paper2Agent generated validated tools for Scanpy's single-cell preprocessing and clustering workflow, allowing an AI agent to execute analysis pipelines from a dataset path.

What is the TISSUE example?

The framework converted TISSUE into an agent for uncertainty-aware spatial-transcriptomics analysis and compared its outputs with those generated by human researchers using the same workflow.

Can scientists use Paper2Agent with their own data?

Potentially yes, when the derived tools support the relevant input format and the scientific method is appropriate for the new dataset.

Can Paper2Agent make new scientific discoveries?

It can help generate and test computational hypotheses, but novel findings still require expert interpretation and independent validation.

Is a Paper2Agent result automatically scientifically correct?

No. Successful tool execution does not guarantee that the source method is scientifically valid or that the interpretation of a new result is correct.

Can Paper2Agent fix a bad paper?

No. It cannot automatically repair flawed methodology, biased data, selective reporting, missing proprietary resources, or an original codebase that never reproduced the claimed findings.

Could Paper2Agent improve reproducibility?

Potentially. It can standardize tool interfaces, expose workflows, execute tutorials, test generated tools, and make computational methods easier for outside researchers to rerun.

Does reproducibility mean a study is correct?

No. A flawed method can be perfectly reproducible. Reproducibility shows that a result or procedure can be repeated; scientific validity asks whether the method and conclusion are sound.

Could paper agents collaborate?

Yes. The researchers demonstrated scenarios in which multiple paper-derived agents could contribute tools to broader analyses. This remains an emerging research direction rather than mature scientific infrastructure.

What is the risk of multiple agents collaborating?

Errors can compound. One incorrect result may become input for another agent, making provenance, validation, and execution logs essential.

Is Paper2Agent open to researchers?

The implementation is publicly available, and the authors provide public code and example MCP servers for systems including AlphaGenome, Scanpy, and TISSUE.

Could paper agents replace scientific papers?

Probably not. Conventional papers provide a stable scholarly record containing evidence, argument, citations, limitations, and interpretation. Agents are more likely to supplement papers than replace them.

What happens if the source paper is corrected?

A mature paper-agent ecosystem would need versioning and a mechanism for propagating corrections to all derived agents. That infrastructure is not yet standardized.

What happens if a paper is retracted?

Ideally, every derived agent should be prominently marked or disabled according to an agreed policy. Otherwise a retracted paper's executable agent could continue circulating.

Could publishers charge for interactive papers?

Yes. Hosted paper agents could become a commercial publishing product, although open community-operated alternatives are also possible.

Could students benefit from Paper2Agent?

Yes. Students could interact directly with methods, change parameters, and observe results. The risk is that easy execution may encourage use without understanding.

Could peer reviewers use paper agents?

Potentially. Reviewers could use them to rerun analyses, identify missing dependencies, and test computational workflows before publication.

Could non-experts misuse Paper2Agent?

Yes. Making advanced scientific software easier to operate can also make it easier to apply incorrectly. Medical, genetic, financial, and other consequential methods need especially clear limits.

Is Paper2Agent an autonomous scientist?

No. It can orchestrate tools and assist with scientific analysis, but humans still need to define important questions, evaluate data quality, assess assumptions, interpret results, and validate discoveries.

Why is versioning important?

Scientific software changes. Libraries, APIs, models, and datasets may be updated after publication, potentially changing or breaking an agent's behavior.

Why is security important for executable papers?

Unlike a PDF, a paper agent can execute code. Future platforms will therefore need protections such as restricted permissions, isolated environments, dependency scanning, and audit logs.

What is the biggest limitation of Paper2Agent?

Its reliability remains strongly dependent on the quality and completeness of the original research package. If code, data, dependencies, or documentation are missing, the agent cannot reliably reconstruct what was never properly preserved.

What is the biggest promise of Paper2Agent?

It could make published scientific methods dramatically easier to reproduce and reuse by turning research software into validated tools accessible through natural language.

What is the most important takeaway?

Paper2Agent suggests that the scientific paper may be evolving from a static document into something more interactive.

But the meaningful future is not simply a world where every paper can chat.

It is a world where research can be read, executed, inspected, reproduced, attributed, versioned, and independently verified.

If Paper2Agent and systems like it can achieve that reliably, scientific publishing may eventually become something researchers do not merely consume.

They may be able to use it directly.

Revlox Magazine Newsletter

Get the latest Revlox stories, cultural essays, and strange discoveries, handpicked for your inbox.

A cleaner edit of the week’s standout reporting, visual culture, historical mysteries, and deeper reads from across the magazine.

By signing up, you agree to the Terms & Conditions and acknowledge the Privacy Policy.

Advertisement

More stories from Revlox Magazine

Read more

AI’s Hidden Water Bill: The 3.4-Trillion-Gallon Data Center Number Is Real—But It Doesn’t Mean 3.4 Trillion Gallons Were “Used Up”

AI’s Hidden Water Bill: The 3.4-Trillion-Gallon Data Center Number Is Real—But It Doesn’t Mean 3.4 Trillion Gallons Were “Used Up”

The modern artificial-intelligence boom does not live entirely in the cloud. It lives in buildings. Enormous buildings. Behind their walls are racks of GPUs and servers consuming electricity continuously, producing heat continuously and depending on physical infrastructure that reaches far beyond the data-center campus itself. Power plants. Transmission

By Imrul

Advertisement

Advertisement

Advertisement