Benchmark 13 of 18, 12 tasks, asked for 2 min 30 s to 3 h 20 min
CORE-Bench v1.1
Reproducing the computational results of a published research paper, given its code capsule, across fields from archaeology to linguistics.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.04× | 72%26 of 36 runs | 34.335 of 36 graded | 34.3%12 of 35 passed | |
| 2 | GPT 5.6 SolCodex | 1.70× | 39%14 of 36 runs | 8.3 | 8.3%3 of 36 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.82× | 6%2 of 36 runs | 0.0 | 0%0 of 36 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. CORE-Bench v1.1 counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Every answer correct, with numbers inside a 95% prediction interval.
107 of 108 runs graded. 100 graded during the run and 7 graded afterwards. Not graded: 1 files not found.
Muse Spark 1.3 has no runs on CORE-Bench v1.1; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every CORE-Bench v1.1 run
Fable's typical run lasted 11 min 15 s at the shortest request and 19 min 37 s at the longest. Astra's lasted 4 min 58 s and 1 h 11 min.
Hover a mark to read it, click to pin it here.
Showing 108 runs.
The 12 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Reproduce a rockshelter chronology and archaeological materials analysis | 10, 40, 150 min | 10.3 min, 41.1 min, 2.5 h | 12.8, 40.5, 15.4 min | 15.8, 17.4, 24.8 min |
| Reproduce article and script-availability results across linguistics | 8, 30, 125 min | 8.4 min, 30.7 min, 2.1 h | 8.6, 33.7, 8.3 min | 6, 7.1, 8.6 min |
| Rebuild the figures for selective exposure to partisan online news | 5, 20, 80 min | 5.6 min, 20.5 min, 1.3 h | 17 min, 20.1 min, 1.3 h | 24.6, 19.4, 23.2 min |
| Run and render CNN short-term temperature forecasting results | 4, 15, 60 min | 4.3 min, 15.4 min, 1 h | 4.9, 15.8, 8.6 min | 12.8, 9.7, 10.2 min |
| Execute the fibrous-media notebook and reproduce its pore-size result | 8, 30, 125 min | 9.1 min, 30.8 min, 2.1 h | 8.5 min, 30.5 min, 2.1 h | 5.1, 11.5, 16.9 min |
| Reproduce journal and university ranking figures | 6, 25, 100 min | 6.3 min, 25.6 min, 1.7 h | 15.7 min, 25.3 min, 1.7 h | 9.7, 16.9, 24.1 min |
| Run a continuous-time intelligent-reflecting-surface propagation model | 4, 15, 60 min | 4.2 min, 15.3 min, 1 h | 14.3, 15.5, 16.9 min | 17.8, 18.6, 29.4 min |
| Reproduce a 20 mm transducer-array characterization | 4, 15, 60 min | 4.2 min, 15.6 min, 1 h | 4.5 min, 15.6 min, 1 h | 5.1, 7.3, 5.8 min |
| Run a quantum measurement and tomography toolbox | 4, 15, 60 min | 4.3 min, 15.6 min, 1 h | 7.6 min, 15.5 min, 1 h | 4.2, 6.6, 11.4 min |
| Execute relative-entropy causal-influence notebooks in parallel | 4, 15, 60 min | 4.2 min, 15.5 min, 1 h | 25.3 min, 17.3 min, 1 h | 3.6, 9.9, 11.7 min |
| Train and evaluate end-to-end neural-network feature selection on MNIST | 13, 50, 200 min | 13.1 min, 50.8 min, 3.4 h | 16.3 min, 3.5 h, 3.4 h | 3.2 h, 6.7 h, 3.1 h |
| Run product-review sentiment mining and reproduce its metrics | 2.5, 10, 40 min | 2.7, 10.8, 44.3 min | 11.6, 14.6, 49 min | 16.5, 11.9, 22.4 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About CORE-Bench v1.1: the prompt, the grader, the source
One CORE-Bench v1.1 task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 150 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by CORE-Bench v1.1's own grader, not a scorer AgentTime invented. The pass rule: Every answer correct, with numbers inside a 95% prediction interval.
- Tasks
- 12
- Asked for
- 2 min 30 s to 3 h 20 min
- Runs here
- 108
- 3 agents
- Graded as
- Pass or fail: Every answer correct, with numbers inside a 95% prediction interval.
- Source
- Nadgir et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Data and citation
Upstream: Nadgir et al., 2026