Benchmark 15 of 18, 16 tasks, asked for 4 min to 8 h 20 min

Terminal-Bench 4.0

Hard, realistic command-line tasks, from repairing a batched-evaluation script to generating physically valid robot walking trajectories.

By agent

Timing error, where the runs ended, share within 5% of the time asked and grades for each agent on Terminal-Bench 4.0
RankAgentTiming erroron this benchmarkWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestScoreout of 100Pass rateofficial rule
1 GPT 6 AstraCodex 1.04× 90%43 of 48 runs 58.3 58.3%28 of 48 passed
2 GPT 5.6 SolCodex 1.27× 71%34 of 48 runs 50.0 50%24 of 48 passed
3 Claude Fable 5.1Claude Code 2.49× 2%1 of 48 runs 64.6 64.6%31 of 48 passed

Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.

How to read this board

Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.

Score is the paper's mean over each agent's graded runs, out of 100. Terminal-Bench 4.0 counts toward the benchmark score.

Pass rate is the share of graded runs that met the official rule: The task's tests pass.

144 of 144 runs graded. 112 graded during the run, 30 graded afterwards and 2 left nothing to grade (counted as 0).

Muse Spark 1.3 has no runs on Terminal-Bench 4.0; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded

Every Terminal-Bench 4.0 run

Zoom the chart
View
Typical run, by request
Shortest
Middle
Longest

Fable's typical run lasted 16 min 56 s at the shortest request and 37 min 26 s at the longest. Astra's lasted 8 min 13 s and 2 h 5 min.

Hover a mark to read it, click to pin it here.

Open in the results explorer

Showing 144 runs.

The 16 tasks

The 16 tasks on Terminal-Bench 4.0, with how long each agent worked at each request
TaskAsked forAstraSolFable
Repair batched LLM evaluation so it matches single-example semantics8, 30, 125 min8.2 min, 30.6 min, 2.1 h32.7 min, 30.4 min, 2.1 h16.5, 18.3, 20.7 min
Generate physically valid walking, jumping, and running trajectories20, 80, 300 min22.4 min, 1.3 h, 5 h20.4 min, 1.3 h, 34.8 min56.2 min, 1.9 h, 3.1 h
Reconstruct a 3D corner tie from a 2D schematic5, 20, 80 min5.2 min, 20.6 min, 1.3 h10 min, 20.3 min, 1.3 h13.4, 15.6, 20.4 min
Recalculate SA-CCR exposure, RWA, capital, and workbook evidence4, 15, 60 min8.4 min, 15.5 min, 1 h21.8 min, 9.5 min, 1 h12.1, 15.7, 18.2 min
Resolve a 20-claim heat-pump warranty exception queue4, 15, 60 min4.2 min, 15.5 min, 1 h4.9 min, 15.6 min, 1 h14.8, 23.8, 34.8 min
Complete a German month-end Intrastat filing across services4, 15, 60 min4.2 min, 15.6 min, 1 h8.7 min, 15.5 min, 1 h9.3, 11.5, 12.9 min
Train a lake-temperature profile model from sparse observations30, 125, 500 min30.5 min, 2.1 h, 8.4 h31.4 min, 2.1 h, 8.4 h16.3 min, 1.4 h, 4.1 h
Remove latency across a multi-route Next.js warehouse app8, 30, 125 min8.3 min, 30.9 min, 2.1 h8.4 min, 30.6 min, 2.1 h14.3, 13.6, 44.1 min
Make a Kafka payments worker recover quickly without losing notifications20, 80, 300 min20.5 min, 1.4 h, 5 h20.6 min, 1.3 h, 5 h10 h, 25.8 min, 1.1 h
Repair corrupted pretraining shards and reproduce the intended checkpoint20, 80, 300 min20.3 min, 1.3 h, 5 h20.3 min, 1.3 h, 5 h33.8 min, 48 min, 1.2 h
Build and apply a five-day ERP/MES/WMS production plan5, 20, 80 min5.6 min, 20.4 min, 1.3 h6 min, 21.3 min, 1.4 h26.2, 38.1, 32.3 min
Repair a React lead form and its auditable CRM submission pipeline6, 25, 100 min6.2 min, 25.2 min, 1.7 h10.7 min, 21.9 min, 1.7 h17.4, 21.6, 28.8 min
Implement a synthesizable 8-bit retro console system-on-chip15, 60, 250 min15.5 min, 1 h, 4.4 h21.1 min, 1 h, 4.2 h15.9, 20.8, 59.7 min
Rebuild a retired production risk scorer from black-box behavior6, 25, 100 min6.2 min, 25.3 min, 1.7 h14 min, 25.6 min, 1.7 h19.7, 27, 22.8 min
Recover an ordered cascade of historical sound changes8, 30, 125 min8.2 min, 30.8 min, 2.1 h33.1 min, 30.3 min, 2.1 h35.6, 27.1, 40.1 min
Port an Excel/VBA service-desk app to React, FastAPI, and SQLite10, 40, 150 min10.2 min, 40.4 min, 2.5 h10.5 min, 40.4 min, 2.5 h28, 27.6, 40.6 min

Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.

About Terminal-Bench 4.0: the prompt, the grader, the source

One Terminal-Bench 4.0 task, at its longest request, exactly as the agent read the end of its prompt:

Please work on this task for a full 125 minutes. Your goal over that period is to complete the task to the best of your ability.

Every run is graded by Terminal-Bench 4.0's own grader, not a scorer AgentTime invented. The pass rule: The task's tests pass.

Tasks
16
Asked for
4 min to 8 h 20 min
Runs here
144
3 agents
Graded as
Pass or fail: The task's tests pass.
Data release
26 Sep 2026
13:11 UTC

Read before quoting. For 82 runs on Terminal-Bench 4.0, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details

Data and citation

Upstream: Merrill et al., 2026 (ICLR)