Method
How we ask, how we time, and how one number per agent comes out of 1,991 runs.
What we ask
AgentTime does not write new tasks. It takes 222 tasks from 18 published benchmarks and adds one sentence to the end of each task's own prompt.
Please work on this task for a full 5 minutes. Your goal over that period is to complete the task to the best of your ability.Only the number and the unit change; every agent gets the same words.
Three requests per task
Every task is asked at three durations, anchored to how long it usually takes: the shortest, about a quarter of the usual time; the middle, close to the usual time; and the longest, about four times the usual time.
An AppWorld task might be asked for 2, 8 and 30 minutes; a ProgramBench task for 12.5 minutes, 50 minutes and 3 hours 20 minutes. Across the suite, requests run from 1 minute to 60 hours.
How we time a run
Each run starts fresh, with the agent's normal tools and no time limit, so it alone decides when to stop. A clock outside the sandbox, invisible to the agent, records how long the session lasts.
Most runs end when the agent decides it is done. A safety cutoff at twice the task's longest request, an error or a provider limit can stop a run instead; its time is then a lower bound, and the chart marks it with a small cap.
Timing error
One number per agent answers a plain question: on a typical run, by what multiple was the agent off? Here is a worked example: asked for 10, 40 and 160 minutes, an agent that worked 20, 40 and 80 misses by 2×, 1× and 2×, for a timing error of about 1.59×.
Timing error for these three runs: 1.59×
- For each run, compute the absolute value of the natural log of worked time divided by requested time.
- Average those values within each benchmark.
- Average the 18 benchmark averages with equal weight.
- Raise e to that power.
Equal weight means PaperBench's 3 tasks count as much as GPQA Diamond's 28, so no single large benchmark decides an agent's number. The 95% intervals resample tasks: GPT 6 Astra 1.18× (1.11 to 1.25), GPT 5.6 Sol 1.77× (1.59 to 1.95), Claude Fable 5.1 2.86× (2.68 to 2.99).
A factor hides direction, which is why every agent's number sits next to its strip and its share of runs within 5% of the time asked.
Within 5% of the time asked
A run is within 5% when it worked between 0.95× and 1.05× of the request, both ends included, for example 38 min to 42 min on a 40 minute request. Below that it is shorter, above it longer. Every board, strip, chart and run on this site uses this window.
Every timed run in the default set counts once, the same runs as each agent's timing error: GPT 6 Astra 666, GPT 5.6 Sol 666, Claude Fable 5.1 659, Muse Spark 1.3 108. The 95% interval resamples tasks with replacement 2,000 times from a fixed seed, so every build prints the same numbers.
Grades
Every run is graded by its own benchmark's grader; AgentTime does not invent a score. 11 of the 18 benchmarks define a pass rule, for example "the hidden tests pass" on DeepSWE v1.1, and the site says Passed or Did not pass there; the others report a score as the benchmark writes it.
The agents
An agent is a model inside a harness: GPT 6 Astra and GPT 5.6 Sol run in Codex; Claude Fable 5.1 runs in Claude Code, each at its maximum reasoning effort. Muse Spark 1.3 is partial and not ranked.
Glossary
- Run
- One agent working one task at one requested time.
- Request
- The duration named in the sentence appended to the task, for example "a full 5 minutes." The agent can return earlier or keep going past it.
- Time worked
- The wall-clock time a run took, measured by a clock outside the agent's sandbox.
- Timing error
- How far a typical run landed from the request, as a multiple. 1.00× is perfect.
- Within 5%
- Worked between 0.95× and 1.05× of the request, both ends included.
- Shorter, longer
- Worked less than 0.95× of the request, or more than 1.05×.
- Ended on its own
- The agent decided by itself that the run was finished.
- Stopped by AgentTime
- A safety cutoff, an error, or a provider limit ended the run, so the time recorded is a lower bound.
- Grade
- The benchmark's own score, kept as it is. "Passed" or "Did not pass" only where the benchmark defines a pass rule.
- Harness
- The tool environment an agent runs inside, for example Codex or Claude Code.
- Partial
- Muse Spark 1.3: runs on 2 of 18 benchmarks, not ranked.