ProgramBench, shortest of three requests

GPT 5.6 Sol on “cweill/gotests”

Longer

1.58× longer

Worked 58% past the request, 7 min 18 s over.

Asked for
12.5 minutes
Worked
19 min 48 s
Grade
Did not pass
Score 62.7 of 100. Graded afterwards.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 5.6 Sol, in Codex
Benchmark
ProgramBench
Asked for
12.5 minutes: “Please work on this task for a full 12.5 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
19 min 48 s (1,188 s)
Ratio
1.58× the request
How it ended
Ended on its own
Grade
Did not pass
It did not pass every hidden test. The grader returned 0.627. Pass rule: Resolved: every hidden test passes.
Score
62.7 of 100
Graded afterwards. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
nht2ht4y5w
Data release
26 Sep 2026
The other requests for GPT 5.6 Sol on this task, and the other agents at this request
Sol on this taskAskedWorkedSplitResult
Shortest (this run)12.5 minutes19 min 48 s1.58× longer
Middle50 minutes50 min 59 sWithin 5%
Longest200 minutes1 h 11 min2.81× shorter
Other agents, shortest request
GPT 6 Astra12.5 minutes12 min 53 sWithin 5%
Claude Fable 5.112.5 minutes1 h 8 min5.43× longer

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Yang et al., 2026