Benchmark 9 of 18, 12 tasks, asked for 1 min 15 s to 1 h 40 min
TUA-Bench
Everyday computer chores done entirely from a terminal: editing documents, managing email, and looking up information on the live web.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.12× | 47%17 of 36 runs | 85.9 | 72.2%26 of 36 passed | |
| 2 | GPT 5.6 SolCodex | 1.31× | 25%9 of 36 runs | 71.235 of 36 graded | 57.1%20 of 35 passed | |
| 3 | Claude Fable 5.1Claude Code | 3.01× | 3%1 of 36 runs | 81.235 of 36 graded | 65.7%23 of 35 passed | |
| Muse Spark 1.3Muse Code. Partial, not ranked | 2.83× | 8%2 of 24 runs | 60.7not the paper's | 41.7%10 of 24 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. TUA-Bench counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Full completion: a reward of 1 on the final state.
130 of 132 runs graded. 126 graded during the run and 4 graded afterwards. Not graded: 2 files not found.
Muse Spark 1.3 is a partial run: shown, never ranked against the paper's three agents, and its grades are the campaign's own, not the paper's. How runs are graded
Every TUA-Bench run
Fable's typical run lasted 3 min 8 s at the shortest request and 7 min 41 s at the longest. Astra's lasted 2 min 47 s and 32 min 59 s.
Hover a mark to read it, click to pin it here.
Showing 108 runs.
The 12 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| APA references review | 2.5, 10, 40 min | 2.8, 10.1, 40.5 min | 3.6, 10.8, 41.1 min | 9.3, 8.9, 9.3 min |
| First-author table | 2.5, 10, 40 min | 2.9, 10.4, 40.6 min | 3.6, 12.4, 41.2 min | 2.1, 2.3, 2.4 min |
| Daily email report | 1.5, 6, 25 min | 1.7, 6.1, 25.2 min | 1.8, 6.1, 25.7 min | 1.2, 2.4, 2.1 min |
| Author homepage bookmarks | 2.5, 10, 40 min | 2.8, 14.3, 40.5 min | 3.7, 11.6, 40.7 min | 2.8, 3.9, 6.1 min |
| Name mountain photos | 1.5, 6, 25 min | 1.8, 6.6, 25.5 min | 2.2, 7.6, 26.3 min | 1.3, 3.1, 2.8 min |
| Merge text document | 1.25, 5, 20 min | 1.8, 5.3, 20.3 min | 3.8, 5.4, 21.2 min | 3.5, 10.6, 9.9 min |
| Place heater for sensors | 4, 15, 60 min | 4.3 min, 15.5 min, 1 h | 8.8 min, 9.6 min, 1 h | 2 h, 22 min, 26.9 min |
| Optimize cold plate | 6, 25, 100 min | 6.5 min, 25.6 min, 1.7 h | 9.9 min, 26 min, 1.7 h | 17.8, 28.1, 35.7 min |
| Extract gym auditorium | 1.25, 5, 20 min | 1.6, 6.1, 20.5 min | 7.1, 6.3, 21.5 min | 21.5, 21, 22.2 min |
| EPW parquet check | 1.5, 6, 25 min | 3.7, 6.4, 25.4 min | 4.3, 6.9, 27.1 min | 7.3, 9.4, 12.3 min |
| Fix MP3 metadata | 1.25, 5, 20 min | 1.5, 5.4, 20.3 min | 1.9, 5.7, 21.1 min | 45 s, 3 min, 1.8 min |
| Extract presenter photos | 2.5, 10, 40 min | 2.8, 10.6, 40.8 min | 2.8, 10.9, 41.6 min | 2.6, 4.5, 5.1 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About TUA-Bench: the prompt, the grader, the source
One TUA-Bench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 40 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by TUA-Bench's own grader, not a scorer AgentTime invented. The pass rule: Full completion: a reward of 1 on the final state.
- Tasks
- 12
- Asked for
- 1 min 15 s to 1 h 40 min
- Runs here
- 132
- 4 agents
- Graded as
- Pass or fail: Full completion: a reward of 1 on the final state.
- Source
- Chen et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 21 runs on TUA-Bench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Chen et al., 2026