Ranked 2 of 3 on timing error

GPT 5.6 Sol

Codex, max reasoning. 666 runs on all 18 benchmarks.

1.77×Timing error95% interval 1.59 to 1.95
39%Within 5%13% shorter, 48% longer
55.6Benchmark scoreOut of 100, 95% interval 50.2 to 60.6

By request

Where GPT 5.6 Sol's runs ended, from a tenth of the time asked to ten times it, for every run and for each of the three requests, with the timing error of each.
RequestTiming errorWhere runs ended
Every timed run666 runs1.77×
Shortest222 runs2.06×
Middle222 runs1.35×
Longest222 runs1.99×

It ran long on short requests (median 1.66× at the shortest) and was within 5% of the time asked on long ones (median 1.01× at the longest). See every run

Every Sol run

Each row is one run, with a mark on the chart. Hover a row to find its mark; click a mark to find its row.

Request
Result

666 runs

GPT 5.6 Sol's runs. Each row links to its run page.
Loading 666 runs.

Across the 18 benchmarks

Timing, and how good the work was. Score is out of 100, except ALE-Bench and YC-Bench, which keep their own units. Pass rate is the share of graded runs that met the benchmark's official rule.

GPT 5.6 Sol across the 18 benchmarks: timing and grades
METR public tasks6 runs
5.72×(few runs)
33%66.7No official pass rule
OSWorld 2.048 runs
2.88×
15%63.320.8%
WildClawBench30 runs
2.72×
13%44.127 of 30 gradedNo official pass rule
Agents' Last Exam36 runs
2.20×
36%66.419.4%
AssistantBench54 runs
2.16×
11%10.63.7%
PostTrainBench v1.16 runs
2.05×(few runs)
50%30.7No official pass rule
PPTArena30 runs
1.89×
30%64.0No official pass rule
YC-Bench9 runs
1.75×(few runs)
22%$0.70Mfinal fundsNo official pass rule
CORE-Bench v1.136 runs
1.70×
39%8.38.3%
ProgramBench42 runs
1.68×
52%69.50%
DeepSWE v1.136 runs
1.50×
56%75.075%
GPQA Diamond84 runs
1.34×
46%94.094%
TUA-Bench36 runs
1.31×
25%71.235 of 36 graded57.1%
PaperBench9 runs
1.31×(few runs)
33%31.8No official pass rule
Terminal-Bench 4.048 runs
1.27×
71%50.050%
Humanity's Last Exam72 runs
1.21×
49%75.075%
AppWorld54 runs
1.13×
28%70.470.4%
Sakana ALE-Bench30 runs
1.12×
73%2518performance ratingNo official pass rule

Highest timing error first. Scores are the paper's, from its Table 5.

662 of 666 runs graded. 473 graded during the run, 156 graded afterwards, 20 judged afterwards and 13 left nothing to grade (counted as 0). Not graded: 4 files not found.

Read before quoting: 2 notes
  1. 17 runs were stopped by AgentTime, so their time is a lower bound. Details
  2. 92 timed runs have an ending not labelled yet. Details

Data and citation