Agents

An agent is a model inside a harness. Each one was asked to work for a set time on the same 222 tasks.

The four agents, ranked by timing error; Muse Spark 1.3 is partial and not ranked.
RankAgentlower is betterWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestout of 100
1 GPT 6 AstraCodex, max reasoning 1.18×1.11 to 1.25 63%419 of 666 runs 61.955.2 to 66.7
2 GPT 5.6 SolCodex, max reasoning 1.77×1.59 to 1.95 39%259 of 666 runs 55.650.2 to 60.6
3 Claude Fable 5.1Claude Code, max reasoning 2.86×2.68 to 2.99 4%27 of 659 runs 54.649.8 to 59.3
Muse Spark 1.3Partial, not ranked: 2 of 18 benchmarks 4.00×no interval 3%3 of 108 runs No benchmark score

Timing error is how far a typical run lands from the time it was asked for; 1.00× is exactly as asked. Where runs ended counts every run from a tenth of the time asked to ten times it; the pale band is within 5% of the time asked. Benchmark score is the paper's mean score out of 100 on 16 benchmarks, over the 604 tasks and requests all three agents have a grade for. Small figures are 95% intervals. How it is measured

The paper also ran Claude Fable 5.1 inside Codex on a 34-task subset, to see whether the harness explains its timing. It did not: Fable still did not track the requested time in either harness. How the agents are set up