Benchmark 15 of 18, 16 tasks, asked for 4 min to 8 h 20 min
Terminal-Bench 4.0
Hard, realistic command-line tasks, from repairing a batched-evaluation script to generating physically valid robot walking trajectories.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.04× | 90%43 of 48 runs | 58.3 | 58.3%28 of 48 passed | |
| 2 | GPT 5.6 SolCodex | 1.27× | 71%34 of 48 runs | 50.0 | 50%24 of 48 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.49× | 2%1 of 48 runs | 64.6 | 64.6%31 of 48 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. Terminal-Bench 4.0 counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: The task's tests pass.
144 of 144 runs graded. 112 graded during the run, 30 graded afterwards and 2 left nothing to grade (counted as 0).
Muse Spark 1.3 has no runs on Terminal-Bench 4.0; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every Terminal-Bench 4.0 run
Fable's typical run lasted 16 min 56 s at the shortest request and 37 min 26 s at the longest. Astra's lasted 8 min 13 s and 2 h 5 min.
Hover a mark to read it, click to pin it here.
Showing 144 runs.
The 16 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Repair batched LLM evaluation so it matches single-example semantics | 8, 30, 125 min | 8.2 min, 30.6 min, 2.1 h | 32.7 min, 30.4 min, 2.1 h | 16.5, 18.3, 20.7 min |
| Generate physically valid walking, jumping, and running trajectories | 20, 80, 300 min | 22.4 min, 1.3 h, 5 h | 20.4 min, 1.3 h, 34.8 min | 56.2 min, 1.9 h, 3.1 h |
| Reconstruct a 3D corner tie from a 2D schematic | 5, 20, 80 min | 5.2 min, 20.6 min, 1.3 h | 10 min, 20.3 min, 1.3 h | 13.4, 15.6, 20.4 min |
| Recalculate SA-CCR exposure, RWA, capital, and workbook evidence | 4, 15, 60 min | 8.4 min, 15.5 min, 1 h | 21.8 min, 9.5 min, 1 h | 12.1, 15.7, 18.2 min |
| Resolve a 20-claim heat-pump warranty exception queue | 4, 15, 60 min | 4.2 min, 15.5 min, 1 h | 4.9 min, 15.6 min, 1 h | 14.8, 23.8, 34.8 min |
| Complete a German month-end Intrastat filing across services | 4, 15, 60 min | 4.2 min, 15.6 min, 1 h | 8.7 min, 15.5 min, 1 h | 9.3, 11.5, 12.9 min |
| Train a lake-temperature profile model from sparse observations | 30, 125, 500 min | 30.5 min, 2.1 h, 8.4 h | 31.4 min, 2.1 h, 8.4 h | 16.3 min, 1.4 h, 4.1 h |
| Remove latency across a multi-route Next.js warehouse app | 8, 30, 125 min | 8.3 min, 30.9 min, 2.1 h | 8.4 min, 30.6 min, 2.1 h | 14.3, 13.6, 44.1 min |
| Make a Kafka payments worker recover quickly without losing notifications | 20, 80, 300 min | 20.5 min, 1.4 h, 5 h | 20.6 min, 1.3 h, 5 h | 10 h, 25.8 min, 1.1 h |
| Repair corrupted pretraining shards and reproduce the intended checkpoint | 20, 80, 300 min | 20.3 min, 1.3 h, 5 h | 20.3 min, 1.3 h, 5 h | 33.8 min, 48 min, 1.2 h |
| Build and apply a five-day ERP/MES/WMS production plan | 5, 20, 80 min | 5.6 min, 20.4 min, 1.3 h | 6 min, 21.3 min, 1.4 h | 26.2, 38.1, 32.3 min |
| Repair a React lead form and its auditable CRM submission pipeline | 6, 25, 100 min | 6.2 min, 25.2 min, 1.7 h | 10.7 min, 21.9 min, 1.7 h | 17.4, 21.6, 28.8 min |
| Implement a synthesizable 8-bit retro console system-on-chip | 15, 60, 250 min | 15.5 min, 1 h, 4.4 h | 21.1 min, 1 h, 4.2 h | 15.9, 20.8, 59.7 min |
| Rebuild a retired production risk scorer from black-box behavior | 6, 25, 100 min | 6.2 min, 25.3 min, 1.7 h | 14 min, 25.6 min, 1.7 h | 19.7, 27, 22.8 min |
| Recover an ordered cascade of historical sound changes | 8, 30, 125 min | 8.2 min, 30.8 min, 2.1 h | 33.1 min, 30.3 min, 2.1 h | 35.6, 27.1, 40.1 min |
| Port an Excel/VBA service-desk app to React, FastAPI, and SQLite | 10, 40, 150 min | 10.2 min, 40.4 min, 2.5 h | 10.5 min, 40.4 min, 2.5 h | 28, 27.6, 40.6 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About Terminal-Bench 4.0: the prompt, the grader, the source
One Terminal-Bench 4.0 task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 125 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by Terminal-Bench 4.0's own grader, not a scorer AgentTime invented. The pass rule: The task's tests pass.
- Tasks
- 16
- Asked for
- 4 min to 8 h 20 min
- Runs here
- 144
- 3 agents
- Graded as
- Pass or fail: The task's tests pass.
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 82 runs on Terminal-Bench 4.0, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Merrill et al., 2026 (ICLR)