Ranked 3 of 3 on timing error
Claude Fable 5.1
Claude Code, max reasoning. 659 runs on all 18 benchmarks.
By request
| Request | Timing error | Where runs ended |
|---|---|---|
| Every timed run659 runs | 2.86× | |
| Shortest221 runs | 3.36× | |
| Middle220 runs | 2.05× | |
| Longest218 runs | 3.26× |
It ran long on short requests (median 1.83× at the shortest) and stopped early on long ones (median 0.31× at the longest). See every run
Every Fable run
Each row is one run, with a mark on the chart. Hover a row to find its mark; click a mark to find its row.
659 runs
| Loading 659 runs. | ||||
No runs match these filters.
Across the 18 benchmarks
Timing, and how good the work was. Score is out of 100, except ALE-Bench and YC-Bench, which keep their own units. Pass rate is the share of graded runs that met the benchmark's official rule.
| PostTrainBench v1.15 runs | 4.98×(few runs) | 0% | 0.03 of 5 graded | No official pass rule |
| METR public tasks6 runs | 4.11×(few runs) | 0% | 62.5 | No official pass rule |
| GPQA Diamond84 runs | 3.76× | 5% | 85.7 | 85.7% |
| WildClawBench30 runs | 3.41× | 10% | 51.8 | No official pass rule |
| TUA-Bench36 runs | 3.01× | 3% | 81.235 of 36 graded | 65.7% |
| AppWorld54 runs | 2.93× | 0% | 75.9 | 75.9% |
| PPTArena30 runs | 2.89× | 3% | 48.7 | No official pass rule |
| CORE-Bench v1.136 runs | 2.82× | 6% | 0.0 | 0% |
| AssistantBench54 runs | 2.75× | 2% | 4.8 | 0% |
| ProgramBench36 runs | 2.75× | 0% | 64.3 | 0% |
| Agents' Last Exam36 runs | 2.68× | 0% | 58.035 of 36 graded | 17.1% |
| PaperBench9 runs | 2.64×(few runs) | 0% | 35.17 of 9 graded | No official pass rule |
| Humanity's Last Exam72 runs | 2.62× | 6% | 87.5 | 87.5% |
| DeepSWE v1.136 runs | 2.57× | 3% | 72.2 | 72.2% |
| Terminal-Bench 4.048 runs | 2.49× | 2% | 64.6 | 64.6% |
| YC-Bench9 runs | 2.43×(few runs) | 0% | $1.11Mfinal funds | No official pass rule |
| OSWorld 2.048 runs | 1.99× | 15% | 67.8 | 27.1% |
| Sakana ALE-Bench30 runs | 1.97× | 7% | 2398performance rating | No official pass rule |
No runs match these filters.
Highest timing error first. Scores are the paper's, from its Table 5.
653 of 659 runs graded. 472 graded during the run, 100 graded afterwards, 46 judged afterwards and 35 left nothing to grade (counted as 0). Not graded: 4 grading…, 1 waiting for its machine and 1 files not found.
Read before quoting: 4 notes
- 1 run on PostTrainBench v1.1 has no finished time in this data release.
- 61 runs were stopped by AgentTime, so their time is a lower bound. Details
- 76 timed runs have an ending not labelled yet. Details
- The paper also ran Claude Fable 5.1 inside Codex on a 34-task subset, to see whether the harness explains its timing. It did not: Claude Fable 5.1 still did not track the requested time in either harness.