TUA-Bench, longest of three requests

GPT 6 Astra on “APA references review”

Within 5%

Within 5%

Worked 101% of the time asked, within 5% of it.

Asked for
40 minutes
Worked
40 min 28 s
Grade
Did not pass
Score 52.5 of 100. Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Benchmark
TUA-Bench
Asked for
40 minutes: “Please work on this task for a full 40 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
40 min 28 s (2,428 s)
Ratio
1.01× the request
How it ended
Ended on its own
Grade
Did not pass
It did not complete the task in full. The grader returned 0.525. Pass rule: Full completion: a reward of 1 on the final state.
Score
52.5 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
wtspodzolb
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest2.5 minutes2 min 45 s1.10× longer
Middle10 minutes10 min 8 sWithin 5%
Longest (this run)40 minutes40 min 28 sWithin 5%
Other agents, longest request
GPT 5.6 Sol40 minutes41 min 5 sWithin 5%
Claude Fable 5.140 minutes9 min 15 s4.32× shorter
Muse Spark 1.340 minutes2 min 51 s14.02× shorter

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Chen et al., 2026