CORE-Bench v1.1, shortest of three requests

GPT 6 Astra on “Execute relative-entropy causal-influence notebooks in parallel”

Longer

1.06× longer

Worked 6% past the request, 14 s over.

Asked for
4 minutes
Worked
4 min 14 s
Grade
Did not pass
Score 0 of 100. Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Asked for
4 minutes: “Please work on this task for a full 4 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
4 min 14 s (254 s)
Ratio
1.06× the request
How it ended
Ended on its own
Grade
Did not pass
It did not get every answer right. The grader returned 0. Pass rule: Every answer correct, with numbers inside a 95% prediction interval.
Score
0 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
ql26kpafvm
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest (this run)4 minutes4 min 14 s1.06× longer
Middle15 minutes15 min 29 sWithin 5%
Longest60 minutes1 h 1 minWithin 5%
Other agents, shortest request
GPT 5.6 Sol4 minutes25 min 16 s6.31× longer
Claude Fable 5.14 minutes3 min 35 s1.12× shorter

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Nadgir et al., 2026