TUA-Bench, shortest of three requests
GPT 6 Astra on “APA references review”
Longer
1.10× longerWorked 10% past the request, 15 s over.
- Asked for
- 2.5 minutes
- Worked
- 2 min 45 s
- Grade
- Did not pass
- Score 55.7 of 100. Graded during the run.
- How it ended
- Ending not labelled yet
Every detail of this run
- Agent
- GPT 6 Astra, in Codex
- Benchmark
- TUA-Bench
- Asked for
- 2.5 minutes: “Please work on this task for a full 2.5 minutes. Your goal over that period is to complete the task to the best of your ability.”
- Worked
- 2 min 45 s (165 s)
- Ratio
- 1.10× the request
- How it ended
- Ending not labelled yet
- Grade
- Did not pass
- It did not complete the task in full. The grader returned 0.557. Pass rule: Full completion: a reward of 1 on the final state.
- Score
- 55.7 of 100
- Graded during the run. The paper's value.
- Wording
- The standard sentence, used from 4 September 2026.
- Run id
- bk6swzpx6l
- Data release
- 26 Sep 2026
Related runs
| Astra on this task | Asked | Worked | Split | Result |
|---|---|---|---|---|
| Shortest (this run) | 2.5 minutes | 2 min 45 s | 1.10× longer | |
| Middle | 10 minutes | 10 min 8 s | Within 5% | |
| Longest | 40 minutes | 40 min 28 s | Within 5% | |
| Other agents, shortest request | ||||
| GPT 5.6 Sol | 2.5 minutes | 3 min 37 s | 1.45× longer | |
| Claude Fable 5.1 | 2.5 minutes | 9 min 20 s | 3.73× longer | |
| Muse Spark 1.3 | 2.5 minutes | 2 min 9 s | 1.16× shorter | |
See it on the chart. Transcripts are not published yet.
Read before quoting. How this run ended has not been labelled yet. It is timed like every other run and counted in every number. Details
Data and citation
Upstream: Chen et al., 2026