TUA-Bench, middle of three requests

GPT 6 Astra on “APA references review”

Within 5%

Within 5%

Worked 101% of the time asked, within 5% of it.

Asked for
10 minutes
Worked
10 min 8 s
Grade
Did not pass
Score 54.4 of 100. Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Benchmark
TUA-Bench
Asked for
10 minutes: “Please work on this task for a full 10 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
10 min 8 s (608 s)
Ratio
1.01× the request
How it ended
Ended on its own
Grade
Did not pass
It did not complete the task in full. The grader returned 0.544. Pass rule: Full completion: a reward of 1 on the final state.
Score
54.4 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
e37ob7ssbw
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest2.5 minutes2 min 45 s1.10× longer
Middle (this run)10 minutes10 min 8 sWithin 5%
Longest40 minutes40 min 28 sWithin 5%
Other agents, middle request
GPT 5.6 Sol10 minutes10 min 49 s1.08× longer
Claude Fable 5.110 minutes8 min 53 s1.12× shorter
Muse Spark 1.310 minutes11 min 8 s1.11× longer

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Chen et al., 2026