AgentTime
Does an AI agent work for as long as it is asked?
222 tasks from 18 benchmarks, each asked at three durations.
Runs within 5% of the time asked
- GPT 6 Astra 63%, 95% interval 59% to 67%,
- GPT 5.6 Sol 39%, 95% interval 35% to 42%,
- Claude Fable 5.1 4%, 95% interval 3% to 6%,
- Muse Spark 1.3Partial, not ranked 3%, 95% interval 0% to 6%,
Duration following
Leaderboard
| Rank | Agent | lower is better | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | out of 100 |
|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex, max reasoning | 1.18×1.11 to 1.25 | 63%419 of 666 runs | 61.955.2 to 66.7 | |
| 2 | GPT 5.6 SolCodex, max reasoning | 1.77×1.59 to 1.95 | 39%259 of 666 runs | 55.650.2 to 60.6 | |
| 3 | Claude Fable 5.1Claude Code, max reasoning | 2.86×2.68 to 2.99 | 4%27 of 659 runs | 54.649.8 to 59.3 | |
| Muse Spark 1.3Partial, not ranked: 2 of 18 benchmarks | 4.00×no interval | 3%3 of 108 runs | No benchmark score |
Small figures are 95% intervals. How it is measured
Every run
Hover a mark to read it, click to pin it. Choose one agent to see its runs as a list.
By benchmark
| Where each agent lands, 1× to 8× | |||||
|---|---|---|---|---|---|
| METR public tasks | 2 | 1.00× | 5.72× | 4.11× | |
| PostTrainBench v1.1 | 2 | 1.01× | 2.05× | 4.98× | |
| WildClawBench | 10 | 1.62× | 2.72× | 3.41× | |
| YC-Bench | 3 | 2.73× | 1.75× | 2.43× | |
| AssistantBench | 18 | 1.68× | 2.16× | 2.75× | |
| GPQA Diamond | 28 | 1.06× | 1.34× | 3.76× | |
| PPTArena | 10 | 1.25× | 1.89× | 2.89× | |
| OSWorld 2.0 | 16 | 1.15× | 2.88× | 1.99× | |
| Agents' Last Exam | 12 | 1.10× | 2.20× | 2.68× | |
| CORE-Bench v1.1 | 12 | 1.04× | 1.70× | 2.82× | |
| ProgramBench | 14 | 1.01× | 1.68× | 2.75× | |
| TUA-Bench | 12 | 1.12× | 1.31× | 3.01× | |
| DeepSWE v1.1 | 12 | 1.07× | 1.50× | 2.57× | |
| AppWorld | 18 | 1.08× | 1.13× | 2.93× | |
| PaperBench | 3 | 1.01× | 1.31× | 2.64× | |
| Humanity's Last Exam | 24 | 1.07× | 1.21× | 2.62× | |
| Terminal-Bench 4.0 | 16 | 1.04× | 1.27× | 2.49× | |
| Sakana ALE-Bench | 10 | 1.02× | 1.12× | 1.97× |
Timing error on each benchmark, highest average first. Grey values rest on fewer than 10 runs. All 18 benchmarks