Benchmark 1 of 18, 18 tasks, asked for 1 min 15 s to 20 min
AssistantBench
Open-ended questions that can only be answered by browsing the live web, such as tracking down a specific number or fact across several pages.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.68× | 11%6 of 54 runs | 3.3 | 0%0 of 54 passed | |
| 2 | GPT 5.6 SolCodex | 2.16× | 11%6 of 54 runs | 10.6 | 3.7%2 of 54 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.75× | 2%1 of 54 runs | 4.8 | 0%0 of 54 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. AssistantBench counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Exact match with the reference answer.
162 of 162 runs graded. 43 graded during the run and 119 graded afterwards.
Muse Spark 1.3 has no runs on AssistantBench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every AssistantBench run
Fable's typical run lasted 3 min 46 s when asked for 1 min 15 s and 13 min 12 s when asked for 20 min. Astra's lasted 1 min 46 s and 18 min 3 s.
Hover a mark to read it, click to pin it here.
Showing 162 runs.
The 18 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Task 1 | 1.25, 5, 20 min | 1.8, 5.7, 18.8 min | 58 s, 2.5 min, 17.1 min | 8.2, 5.5, 15.1 min |
| Task 2 | 1.25, 5, 20 min | 1.7, 3, 6.8 min | 2.8, 3, 5.5 min | 10.7, 32.2, 40 min |
| Task 3 | 1.25, 5, 20 min | 1.7, 5.3, 5.8 min | 6.2, 3.7, 18.6 min | 3.1, 5.3, 8.1 min |
| Task 4 | 1.25, 5, 20 min | 1.5, 1.7, 2.1 min | 2.9, 5.2, 2.5 min | 1, 2, 2.1 min |
| Task 5 | 1.25, 5, 20 min | 1.9, 6.2, 9.6 min | 3.6, 2.2, 12 min | 2.1, 6.6, 40 min |
| Task 6 | 1.25, 5, 20 min | 2.4, 5.6, 18.9 min | 7.9, 3.8, 12.6 min | 40, 40, 40 min |
| Task 7 | 1.25, 5, 20 min | 1.7, 8.3, 20.4 min | 5.8, 3.8, 18.7 min | 11.1, 12.1, 11.8 min |
| Task 8 | 1.25, 5, 20 min | 4.6, 14.6, 19 min | 3, 31.4, 37.8 min | 6.1, 8.4, 14.6 min |
| Task 9 | 1.25, 5, 20 min | 29 s, 47 s, 53 s | 3.1, 2.2, 5.4 min | 48 s, 1.3 min, 1.5 min |
| Task 10 | 1.25, 5, 20 min | 2.1, 5.5, 17.5 min | 7.1, 5.1, 10.7 min | 7.3, 14.1, 40 min |
| Task 11 | 1.25, 5, 20 min | 1.8, 5.7, 20.8 min | 1.9, 5.1, 18.7 min | 2.1, 1.8, 2.6 min |
| Task 12 | 1.25, 5, 20 min | 1.7, 6.5, 20.9 min | 9.1, 5.3, 19.9 min | 2.3, 21.8, 15.7 min |
| Task 13 | 1.25, 5, 20 min | 1.5, 7.1, 18.6 min | 4.8, 4.4, 4.3 min | 7.1, 10.2, 18.1 min |
| Task 14 | 1.25, 5, 20 min | 1.9, 5.9, 6.9 min | 15.1, 12.2, 13.5 min | 1.3, 5.8, 4.7 min |
| Task 15 | 1.25, 5, 20 min | 1.8, 5.6, 20.3 min | 3.2, 4.3, 18.7 min | 4.4, 5.5, 10.9 min |
| Task 16 | 1.25, 5, 20 min | 3, 5.8, 4.9 min | 8.5, 2.5, 19.3 min | 2.1, 6.2, 4.4 min |
| Task 17 | 1.25, 5, 20 min | 1.4, 5.5, 20.5 min | 3, 4.1, 20.6 min | 1.5, 3.5, 5.9 min |
| Task 18 | 1.25, 5, 20 min | 4.3, 6.3, 13.8 min | 9.7, 9.7, 14.2 min | 40, 8.2, 40 min |
No runs match these filters.
Every task is an AssistantBench validation task. Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About AssistantBench: the prompt, the grader, the source
One AssistantBench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 20 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by AssistantBench's own grader, not a scorer AgentTime invented. The pass rule: Exact match with the reference answer.
- Tasks
- 18
- Asked for
- 1 min 15 s to 20 min
- Runs here
- 162
- 3 agents
- Graded as
- Pass or fail: Exact match with the reference answer.
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 8 runs on AssistantBench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Yoran et al., 2024 (EMNLP)