Agents
An agent is a model inside a harness. Each one was asked to work for a set time on the same 222 tasks.
| Rank | Agent | lower is better | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | out of 100 |
|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex, max reasoning | 1.18×1.11 to 1.25 | 63%419 of 666 runs | 61.955.2 to 66.7 | |
| 2 | GPT 5.6 SolCodex, max reasoning | 1.77×1.59 to 1.95 | 39%259 of 666 runs | 55.650.2 to 60.6 | |
| 3 | Claude Fable 5.1Claude Code, max reasoning | 2.86×2.68 to 2.99 | 4%27 of 659 runs | 54.649.8 to 59.3 | |
| Muse Spark 1.3Partial, not ranked: 2 of 18 benchmarks | 4.00×no interval | 3%3 of 108 runs | No benchmark score |
Timing error is how far a typical run lands from the time it was asked for; 1.00× is exactly as asked. Where runs ended counts every run from a tenth of the time asked to ten times it; the pale band is within 5% of the time asked. Benchmark score is the paper's mean score out of 100 on 16 benchmarks, over the 604 tasks and requests all three agents have a grade for. Small figures are 95% intervals. How it is measured
The paper also ran Claude Fable 5.1 inside Codex on a 34-task subset, to see whether the harness explains its timing. It did not: Fable still did not track the requested time in either harness. How the agents are set up