Ranked 1 of 3 on timing error

GPT 6 Astra

Codex, max reasoning. 666 runs on all 18 benchmarks.

1.18×Timing error95% interval 1.11 to 1.25
63%Within 5%3% shorter, 34% longer
61.9Benchmark scoreOut of 100, 95% interval 55.2 to 66.7

By request

Where GPT 6 Astra's runs ended, from a tenth of the time asked to ten times it, for every run and for each of the three requests, with the timing error of each.
RequestTiming errorWhere runs ended
Every timed run666 runs1.18×
Shortest222 runs1.22×
Middle222 runs1.17×
Longest222 runs1.15×

It ran long on short requests (median 1.12× at the shortest) and was within 5% of the time asked on long ones (median 1.01× at the longest). See every run

Every Astra run

Each row is one run, with a mark on the chart. Hover a row to find its mark; click a mark to find its row.

Request
Result

666 runs

GPT 6 Astra's runs. Each row links to its run page.
Loading 666 runs.

Across the 18 benchmarks

Timing, and how good the work was. Score is out of 100, except ALE-Bench and YC-Bench, which keep their own units. Pass rate is the share of graded runs that met the benchmark's official rule.

GPT 6 Astra across the 18 benchmarks: timing and grades
YC-Bench9 runs
2.73×(few runs)
22%$0.67Mfinal fundsNo official pass rule
AssistantBench54 runs
1.68×
11%3.30%
WildClawBench30 runs
1.62×
37%51.427 of 30 gradedNo official pass rule
PPTArena30 runs
1.25×
47%64.0No official pass rule
OSWorld 2.048 runs
1.15×
63%73.725%
TUA-Bench36 runs
1.12×
47%85.972.2%
Agents' Last Exam36 runs
1.10×
75%66.325%
AppWorld54 runs
1.08×
48%61.161.1%
DeepSWE v1.136 runs
1.07×
67%77.877.8%
Humanity's Last Exam72 runs
1.07×
64%77.877.8%
GPQA Diamond84 runs
1.06×
65%96.496.4%
CORE-Bench v1.136 runs
1.04×
72%34.335 of 36 graded34.3%
Terminal-Bench 4.048 runs
1.04×
90%58.358.3%
Sakana ALE-Bench30 runs
1.02×
97%3067performance rating28 of 30 gradedNo official pass rule
PaperBench9 runs
1.01×(few runs)
100%40.14 of 9 gradedNo official pass rule
ProgramBench42 runs
1.01×
100%74.00%
PostTrainBench v1.16 runs
1.01×(few runs)
100%75.94 of 6 gradedNo official pass rule
METR public tasks6 runs
1.00×(few runs)
100%60.05 of 6 gradedNo official pass rule

Highest timing error first. Scores are the paper's, from its Table 5.

652 of 666 runs graded. 462 graded during the run, 174 graded afterwards, 8 judged afterwards and 8 left nothing to grade (counted as 0). Not graded: 7 grading…, 6 files not found and 1 grader refused.

Read before quoting: 2 notes
  1. 15 runs were stopped by AgentTime, so their time is a lower bound. Details
  2. 91 timed runs have an ending not labelled yet. Details

Data and citation