Benchmark 6 of 18, 3 tasks, asked for 5 min to 3 h 20 min
PaperBench
Reproducing a machine learning research paper from scratch, using only the paper itself and its addendum, no original code.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.01× | 100%9 of 9 runs | 40.14 of 9 graded | No official pass rule | |
| 2 | GPT 5.6 SolCodex | 1.31× | 33%3 of 9 runs | 31.8 | No official pass rule | |
| 3 | Claude Fable 5.1Claude Code | 2.64× | 0%0 of 9 runs | 35.17 of 9 graded | No official pass rule |
Score is the paper's mean over each agent's graded runs, out of 100. PaperBench has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. PaperBench counts toward the benchmark score.
This benchmark has no official pass rule, so there is no pass rate.
20 of 27 runs graded. 19 judged afterwards and 1 left nothing to grade (counted as 0). Not graded: 7 grading….
Muse Spark 1.3 has no runs on PaperBench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every PaperBench run
Fable's typical run lasted 5 min 18 s at the shortest request and 42 min 3 s at the longest. Astra's lasted 10 min 9 s and 2 h 31 min.
Hover a mark to read it, click to pin it here.
Showing 27 runs.
The 3 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Bridging Data Gaps | 13, 50, 200 min | 12.7 min, 50.6 min, 3.3 h | 6.2 min, 51.3 min, 3.3 h | 5.3, 15.2, 45.5 min |
| FRE | 10, 40, 150 min | 10.2 min, 41.6 min, 2.5 h | 6.4 min, 27.6 min, 2.5 h | 5, 12.6, 42.1 min |
| BAM | 5, 20, 80 min | 5.1 min, 20.2 min, 1.3 h | 8.6 min, 24.1 min, 1.5 h | 6.2, 11.3, 22 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About PaperBench: the prompt, the grader, the source
One PaperBench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 200 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by PaperBench's own grader, not a scorer AgentTime invented. It reports a score: Replication score: a weighted rubric of pass/fail leaf checks, 0 to 1, after a full reproduction attempt.
- Tasks
- 3
- Asked for
- 5 min to 3 h 20 min
- Runs here
- 27
- 3 agents
- Graded as
- Replication score: a weighted rubric of pass/fail leaf checks, 0 to 1, after a full reproduction attempt.
- Source
- Starace et al., 2025
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting: 2 notes
- PaperBench has 3 tasks, so one run can move its numbers by a lot.
- For 9 runs on PaperBench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Starace et al., 2025