Method

How we ask, how we time, and how one number per agent comes out of 1,991 runs.

What we ask

AgentTime does not write new tasks. It takes 222 tasks from 18 published benchmarks and adds one sentence to the end of each task's own prompt.

Please work on this task for a full 5 minutes. Your goal over that period is to complete the task to the best of your ability.

Only the number and the unit change; every agent gets the same words.

Three requests per task

Every task is asked at three durations, anchored to how long it usually takes: the shortest, about a quarter of the usual time; the middle, close to the usual time; and the longest, about four times the usual time.

Shortest1/4×Middle1×Longest4×

How we time a run

Each run starts fresh, with the agent's normal tools and no time limit, so it alone decides when to stop. A clock outside the sandbox, invisible to the agent, records how long the session lasts.

Most runs end when the agent decides it is done. A safety cutoff at twice the task's longest request, an error or a provider limit can stop a run instead; its time is then a lower bound, and the chart marks it with a small cap.

Timing error

One number per agent answers a plain question: on a typical run, by what multiple was the agent off? Here is a worked example: asked for 10, 40 and 160 minutes, an agent that worked 20, 40 and 80 misses by 2×, 1× and 2×, for a timing error of about 1.59×.

Left of 1× the run ended early, right of it the run went long. The pale band is within 5% of the time asked.

Timing error for these three runs: 1.59×

Within 5% of the time asked

A run is within 5% when it worked between 0.95× and 1.05× of the request, both ends included, for example 38 min to 42 min on a 40 minute request. Below that it is shorter, above it longer. Every board, strip, chart and run on this site uses this window.

Grades

Every run is graded by its own benchmark's grader; AgentTime does not invent a score. 11 of the 18 benchmarks define a pass rule, for example "the hidden tests pass" on DeepSWE v1.1, and the site says Passed or Did not pass there; the others report a score as the benchmark writes it.

The agents

An agent is a model inside a harness: GPT 6 Astra and GPT 5.6 Sol run in Codex; Claude Fable 5.1 runs in Claude Code, each at its maximum reasoning effort. Muse Spark 1.3 is partial and not ranked.

Glossary

Back to top