CORE-Bench v1.1, middle of three requests

GPT 6 Astra on “Run product-review sentiment mining and reproduce its metrics”

Longer

1.07× longer

Worked 7% past the request, 45 s over.

Asked for
10 minutes
Worked
10 min 45 s
Grade
Passed
Score 100 of 100. Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Asked for
10 minutes: “Please work on this task for a full 10 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
10 min 45 s (645 s)
Ratio
1.07× the request
How it ended
Ended on its own
Grade
Passed
It got every answer right. The grader returned 1. Pass rule: Every answer correct, with numbers inside a 95% prediction interval.
Score
100 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
nvs32uwnrs
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest2.5 minutes2 min 44 s1.09× longer
Middle (this run)10 minutes10 min 45 s1.07× longer
Longest40 minutes44 min 15 s1.11× longer
Other agents, middle request
GPT 5.6 Sol10 minutes14 min 36 s1.46× longer
Claude Fable 5.110 minutes11 min 56 s1.19× longer

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Nadgir et al., 2026