TUA-Bench, longest of three requests

GPT 6 Astra on “EPW parquet check”

Within 5%

Within 5%

Worked 102% of the time asked, within 5% of it.

Asked for
25 minutes
Worked
25 min 23 s
Grade
Did not pass
Score 66 of 100. Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Benchmark
TUA-Bench
Asked for
25 minutes: “Please work on this task for a full 25 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
25 min 23 s (1,523 s)
Ratio
1.02× the request
How it ended
Ended on its own
Grade
Did not pass
It did not complete the task in full. The grader returned 0.66. Pass rule: Full completion: a reward of 1 on the final state.
Score
66 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
isvhlr4bac
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest1.5 minutes3 min 42 s2.47× longer
Middle6 minutes6 min 23 s1.06× longer
Longest (this run)25 minutes25 min 23 sWithin 5%
Other agents, longest request
GPT 5.6 Sol25 minutes27 min 6 s1.08× longer
Claude Fable 5.125 minutes12 min 16 s2.04× shorter
Muse Spark 1.325 minutes5 min 48 s4.31× shorter

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Chen et al., 2026