Terminal-Bench 4.0, middle of three requests

Claude Fable 5.1 on “Generate physically valid walking, jumping, and running trajectories”

Longer

1.40× longer

Worked 40% past the request, 32 min 5 s over.

Asked for
80 minutes
Worked
1 h 52 min
Grade
Did not pass
Score 0 of 100. Nothing to grade, counted as 0.
How it ended
Stopped by an error
Every detail of this run
Agent
Claude Fable 5.1, in Claude Code
Asked for
80 minutes: “Please work on this task for a full 80 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
1 h 52 min (6,725 s)
Ratio
1.40× the request
How it ended
Stopped by an error
Grade
Did not pass
It did not pass the task's tests. The grader returned 0. Pass rule: The task's tests pass.
Score
0 of 100
Nothing to grade, counted as 0. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
ua5e33hpd7
Data release
26 Sep 2026
The other requests for Claude Fable 5.1 on this task, and the other agents at this request
Fable on this taskAskedWorkedSplitResult
Shortest20 minutes56 min 9 s2.81× longer
Middle (this run)80 minutes1 h 52 min1.40× longer
Longest300 minutes3 h 6 min1.61× shorter
Other agents, middle request
GPT 6 Astra80 minutes1 h 20 minWithin 5%
GPT 5.6 Sol80 minutes1 h 21 minWithin 5%

See it on the chart. Transcripts are not published yet.

Read before quoting. This run was stopped by an error, so its time is a lower bound: the agent might have kept going. Details

Data and citation

Upstream: Merrill et al., 2026 (ICLR)