Benchmark 17 of 18, 2 tasks, asked for 2 h to 60 h
METR public tasks
Two of METR's public autonomy tasks: building an AI to play a novel board game, and building fraud-style duplicate-payment detection across currencies and time zones.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.00× | 100%6 of 6 runs | 60.05 of 6 graded | No official pass rule | |
| 2 | Claude Fable 5.1Claude Code | 4.11× | 0%0 of 6 runs | 62.5 | No official pass rule | |
| 3 | GPT 5.6 SolCodex | 5.72× | 33%2 of 6 runs | 66.7 | No official pass rule |
Score is the paper's mean over each agent's graded runs, out of 100. METR public tasks has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. METR public tasks counts toward the benchmark score.
This benchmark has no official pass rule, so there is no pass rate.
17 of 18 runs graded. 6 graded during the run, 8 graded afterwards and 3 left nothing to grade (counted as 0). Not graded: 1 grader refused.
Muse Spark 1.3 has no runs on METR public tasks; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every METR public tasks run
Sol's typical run lasted 3 h 4 min at the shortest request and 3 h 19 min at the longest. Astra's lasted 3 h 1 min and 45 h 5 min.
Hover a mark to read it, click to pin it here.
Showing 18 runs.
The 2 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Build an efficient AI for the novel Cowthello board game | 2 h, 8 h, 30 h | 2 h, 8 h, 30 h | 2.1 h, 8 h, 6.1 h | 37.5 min, 3.4 h, 8.4 h |
| Implement duplicate-payment detection across currencies and time zones | 4 h, 15 h, 60 h | 4 h, 15 h, 60.1 h | 4 h, 15.6 min, 31.2 min | 1.3 h, 4.8 h, 3.2 h |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About METR public tasks: the prompt, the grader, the source
One METR public tasks task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 30 hours. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by METR public tasks's own grader, not a scorer AgentTime invented. It reports a score: Cowthello: win rate against 3 bots, divided by a 70% target and capped at 1. Complex Payments: share of held-out feature tests passed.
- Tasks
- 2
- Asked for
- 2 h to 60 h
- Runs here
- 18
- 3 agents
- Graded as
- Cowthello: win rate against 3 bots, divided by a 70% target and capped at 1. Complex Payments: share of held-out feature tests passed.
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting: 2 notes
- METR public tasks has 2 tasks, so one run can move its numbers by a lot.
- For 4 runs on METR public tasks, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Kinniment et al., 2024 (METR)