AgentTime

Does an AI agent work for as long as it is asked?

222 tasks from 18 benchmarks, each asked at three durations.

Runs within 5% of the time asked

  1. GPT 6 Astra 63%, 95% interval 59% to 67%, 419 of 666 runs
  2. GPT 5.6 Sol 39%, 95% interval 35% to 42%, 259 of 666 runs
  3. Claude Fable 5.1 4%, 95% interval 3% to 6%, 27 of 659 runs
  • Muse Spark 1.3Partial, not ranked 3%, 95% interval 0% to 6%, 3 of 108 runs
Higher is better. Lines are 95% intervals. How it is counted

Duration following

Leaderboard

Agents ranked by timing error, lower is better. 1.00× means every run worked exactly as long as asked. Benchmark score is out of 100.
RankAgentlower is betterWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestout of 100
1 GPT 6 AstraCodex, max reasoning 1.18×1.11 to 1.25 63%419 of 666 runs 61.955.2 to 66.7
2 GPT 5.6 SolCodex, max reasoning 1.77×1.59 to 1.95 39%259 of 666 runs 55.650.2 to 60.6
3 Claude Fable 5.1Claude Code, max reasoning 2.86×2.68 to 2.99 4%27 of 659 runs 54.649.8 to 59.3
Muse Spark 1.3Partial, not ranked: 2 of 18 benchmarks 4.00×no interval 3%3 of 108 runs No benchmark score

Small figures are 95% intervals. How it is measured

Every run

Hover a mark to read it, click to pin it. Choose one agent to see its runs as a list.

Open the results explorer

By benchmark

Timing error on each of the 18 benchmarks for the three agents in the paper, highest average first.
Where each agent lands, 1× to 8×
METR public tasks 2 1.00×5.72×4.11×
PostTrainBench v1.1 2 1.01×2.05×4.98×
WildClawBench 10 1.62×2.72×3.41×
YC-Bench 3 2.73×1.75×2.43×
AssistantBench 18 1.68×2.16×2.75×
GPQA Diamond 28 1.06×1.34×3.76×
PPTArena 10 1.25×1.89×2.89×
OSWorld 2.0 16 1.15×2.88×1.99×
Agents' Last Exam 12 1.10×2.20×2.68×
CORE-Bench v1.1 12 1.04×1.70×2.82×
ProgramBench 14 1.01×1.68×2.75×
TUA-Bench 12 1.12×1.31×3.01×
DeepSWE v1.1 12 1.07×1.50×2.57×
AppWorld 18 1.08×1.13×2.93×
PaperBench 3 1.01×1.31×2.64×
Humanity's Last Exam 24 1.07×1.21×2.62×
Terminal-Bench 4.0 16 1.04×1.27×2.49×
Sakana ALE-Bench 10 1.02×1.12×1.97×

Timing error on each benchmark, highest average first. Grey values rest on fewer than 10 runs. All 18 benchmarks