WildClawBench, shortest of three requests

GPT 6 Astra on “Triage support messages, investigate hidden context, route issues, and draft replies”

Longer

3.12× longer

Worked 212% past the request, 2 min 7 s over.

Asked for
1 minute
Worked
3 min 7 s
Grade
Score 77.5 of 100
Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Benchmark
WildClawBench
Asked for
1 minute: “Please work on this task for a full 1 minute. Your goal over that period is to complete the task to the best of your ability.”
Worked
3 min 7 s (187 s)
Ratio
3.12× the request
How it ended
Ended on its own
Grade
0.775
Task grader score, 0 to 1 (a model judge only where the task calls for one). No official pass rule.
Score
77.5 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
3vqm77v7ht
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest (this run)1 minute3 min 7 s3.12× longer
Middle4 minutes4 min 23 s1.10× longer
Longest15 minutes15 min 31 sWithin 5%
Other agents, shortest request
GPT 5.6 Sol1 minute7 min 30 s7.50× longer
Claude Fable 5.11 minute10 min 3 s10.05× longer

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Ding et al., 2026