Ranked 3 of 3 on timing error

Claude Fable 5.1

Claude Code, max reasoning. 659 runs on all 18 benchmarks.

2.86×Timing error95% interval 2.68 to 2.99
4%Within 5%55% shorter, 41% longer
54.6Benchmark scoreOut of 100, 95% interval 49.8 to 59.3

By request

Where Claude Fable 5.1's runs ended, from a tenth of the time asked to ten times it, for every run and for each of the three requests, with the timing error of each.
RequestTiming errorWhere runs ended
Every timed run659 runs2.86×
Shortest221 runs3.36×
Middle220 runs2.05×
Longest218 runs3.26×

It ran long on short requests (median 1.83× at the shortest) and stopped early on long ones (median 0.31× at the longest). See every run

Every Fable run

Each row is one run, with a mark on the chart. Hover a row to find its mark; click a mark to find its row.

Request
Result

659 runs

Claude Fable 5.1's runs. Each row links to its run page.
Loading 659 runs.

Across the 18 benchmarks

Timing, and how good the work was. Score is out of 100, except ALE-Bench and YC-Bench, which keep their own units. Pass rate is the share of graded runs that met the benchmark's official rule.

Claude Fable 5.1 across the 18 benchmarks: timing and grades
PostTrainBench v1.15 runs
4.98×(few runs)
0%0.03 of 5 gradedNo official pass rule
METR public tasks6 runs
4.11×(few runs)
0%62.5No official pass rule
GPQA Diamond84 runs
3.76×
5%85.785.7%
WildClawBench30 runs
3.41×
10%51.8No official pass rule
TUA-Bench36 runs
3.01×
3%81.235 of 36 graded65.7%
AppWorld54 runs
2.93×
0%75.975.9%
PPTArena30 runs
2.89×
3%48.7No official pass rule
CORE-Bench v1.136 runs
2.82×
6%0.00%
AssistantBench54 runs
2.75×
2%4.80%
ProgramBench36 runs
2.75×
0%64.30%
Agents' Last Exam36 runs
2.68×
0%58.035 of 36 graded17.1%
PaperBench9 runs
2.64×(few runs)
0%35.17 of 9 gradedNo official pass rule
Humanity's Last Exam72 runs
2.62×
6%87.587.5%
DeepSWE v1.136 runs
2.57×
3%72.272.2%
Terminal-Bench 4.048 runs
2.49×
2%64.664.6%
YC-Bench9 runs
2.43×(few runs)
0%$1.11Mfinal fundsNo official pass rule
OSWorld 2.048 runs
1.99×
15%67.827.1%
Sakana ALE-Bench30 runs
1.97×
7%2398performance ratingNo official pass rule

Highest timing error first. Scores are the paper's, from its Table 5.

653 of 659 runs graded. 472 graded during the run, 100 graded afterwards, 46 judged afterwards and 35 left nothing to grade (counted as 0). Not graded: 4 grading…, 1 waiting for its machine and 1 files not found.

Read before quoting: 4 notes
  1. 1 run on PostTrainBench v1.1 has no finished time in this data release.
  2. 61 runs were stopped by AgentTime, so their time is a lower bound. Details
  3. 76 timed runs have an ending not labelled yet. Details
  4. The paper also ran Claude Fable 5.1 inside Codex on a 34-task subset, to see whether the harness explains its timing. It did not: Claude Fable 5.1 still did not track the requested time in either harness.

Data and citation