Data
Every number on this site comes from these files.
Release 2026-09-26-1311Z. Original data: 13:11 UTC on 26 Sep 2026.
Files
| File | Rows | Size | SHA-256 (first 12) | What it holds |
|---|---|---|---|---|
| agenttime-runs.csv | 2,105 | 612.8 KB | 506a17107418 | One row per public run. |
| agenttime-runs.json | 2,105 | 1.7 MB | 9a6de3a6db21 | The same runs, with the grade nested as one object. |
| agenttime-tasks.csv | 223 | 35.7 KB | cbf769526f8e | One row per task: title, benchmark and the three requests. |
| agenttime-benchmarks.csv | 18 | 4.6 KB | db24cfa0ac51 | One row per benchmark: public facts and per-agent timing error. |
| agenttime-agents.csv | 4 | 1 KB | 64695d4ef918 | One row per agent: timing error, shares and endings. |
| schema.json | 25.5 KB | 135f9aeda22d | Field definitions for every file in this release. | |
| README.txt | 1.8 KB | 2f0e015eeaae | What each file holds, in plain words. | |
| SHA256SUMS | 7 | 590 B | 6257b35ad271 | A SHA-256 checksum for every other file in this folder. |
No runs match these filters.
8 files, 2.4 MB in all.
Fields
Every column, its type and its meaning. The full list, machine readable, is in schema.json.
The JSON file nests grade_display, grade_value and passed as one grade object; release_id sits once at the top of the file instead of on every run.
| Field | Type | Meaning |
|---|---|---|
| release_id | string | The data release this run belongs to. |
| run_id | string | Opaque id for one run: ten lowercase letters and digits. |
| agent | string | The agent’s id, for example claude-fable-5-1. |
| agent_name | string | The agent’s full name, for example Claude Fable 5.1. |
| harness | string | The coding harness the agent ran inside, for example Claude Code. |
| benchmark | string | The benchmark’s slug, for example deepswe. |
| benchmark_name | string | The benchmark’s name, for example DeepSWE. |
| task | string | The task’s slug, unique within its benchmark. |
| task_raw_id | string | The benchmark’s own id for the task. |
| task_title | string | The task’s title. |
| request | string | Which of the task’s three requests this run answers: shortest, middle or longest. |
| requested_seconds | number | How many seconds the appended sentence asked the agent to work. |
| worked_seconds | number or empty | Recorded elapsed seconds. Normally measured outside the sandbox; a recovered CLI-duration estimate is identified by timing_basis and timing_estimated. Empty when no duration is available. |
| ratio | number or empty | worked_seconds divided by requested_seconds, to four decimals. Empty when no duration is available. |
| verdict | string or empty | The release’s original label: early, on_time or late, with on_time a ratio between 0.8 and 1.25, both included. The site does not use it. Empty when no duration is available. |
| within_5pct | boolean or empty | True when the ratio is between 0.95 and 1.05, both included: the site’s within 5% of the time asked, which every page uses. Empty when no duration is available. |
| ending | string | How the run ended: own (the agent stopped by itself), harness (a safety cutoff, an error or a provider limit ended it), or unlabelled (not labelled yet). |
| ending_detail | string or empty | When ending is harness: safety_cutoff, error or provider_limit. Empty otherwise. |
| archived | boolean | True when the run has no recorded clock tool call. Only possible on GPQA Diamond and Humanity’s Last Exam. Archived runs are still counted. |
| refusal_left_out | boolean | True for the six Claude Fable 5.1 ProgramBench runs that refused within 30 seconds. Left out of the default set. |
| retired_question | boolean | True for runs on the Humanity’s Last Exam question replaced on 25 September 2026. |
| clock_tool | string | provided, not_recorded or not_applicable. Only GPQA Diamond and Humanity’s Last Exam runs carry a value other than not_applicable. |
| grade_display | string | The benchmark’s own score, written out, or "Not graded". |
| grade_value | number or empty | The score as one number, when it reduces to one. Empty otherwise. |
| passed | boolean or empty | True or false where the benchmark defines a pass rule. Empty where it does not, or where the run is not graded. |
| prompt_version | string | The wording version of the appended request sentence. |
| in_default | boolean | False only for the six left-out refusals. Timing statistics require recorded timing, including labelled recovered estimates; paper score aggregates retain the paper population. |
| completion_status | string or empty | completed for a separately admitted checkpoint result. Empty for historical rows. |
| completion_basis | string or empty | retained_checkpoint when a saved, graded checkpoint supplies the completed result. |
| checkpoint_date | string or empty | Date the admitted checkpoint was saved, not the date of a later attempt. |
| graded_at | string or empty | When the recovery evaluation completed, in UTC. |
| completion_note | string or empty | Public recovery provenance, timing limits and grading conditions. |
| grade_receipt_sha256 | string or empty | SHA-256 of the retained grading receipt supporting this completion. |
| timing_basis | string or empty | native_cli_duration for elapsed time from the identified timing attempt’s CLI terminal event. The grade may come from a separate attempt. |
| timing_estimated | boolean or empty | True when a recovered CLI duration is included as an estimate because the full-process end receipt is absent. |
| timing_receipt_sha256 | string or empty | SHA-256 of the separate receipt documenting the recovered timing evidence. |
| grade_attempt | string or empty | Public date and revision of the attempt that produced the graded checkpoint. |
| timing_attempt | string or empty | Public date and revision of the attempt supplying the elapsed time. |
| timing_same_attempt_as_grade | boolean or empty | False when the displayed time and grade come from separate attempts at the same task and time request. They must not be interpreted as a measured time-and-score pair. |
| timing_supersedes_receipt_sha256 | string or empty | SHA-256 of a previous website timing receipt explicitly corrected by this admission. |
No runs match these filters.
| Field | Type | Meaning |
|---|---|---|
| release_id | string | The data release this row belongs to. |
| benchmark | string | The benchmark’s slug. |
| task | string | The task’s slug, unique within its benchmark. |
| task_raw_id | string | The benchmark’s own id for the task. |
| title | string | The task’s title. |
| retired | boolean | True for the one Humanity’s Last Exam question replaced on 25 September 2026. |
| shortest_seconds | number | The task’s shortest request, in seconds. |
| shortest_words | string | The shortest request, in the words the prompt used, for example "8 minutes". |
| middle_seconds | number | The task’s middle request, in seconds. |
| middle_words | string | The middle request, in words. |
| longest_seconds | number | The task’s longest request, in seconds. |
| longest_words | string | The longest request, in words. |
No runs match these filters.
| Field | Type | Meaning |
|---|---|---|
| release_id | string | The data release this row belongs to. |
| benchmark | string | The benchmark’s slug. |
| benchmark_name | string | The benchmark’s name. |
| task_count | number | How many tasks the benchmark contributes to the roster. |
| requests_min_seconds | number | The shortest request asked anywhere on this benchmark, in seconds. |
| requests_max_seconds | number | The longest request asked anywhere on this benchmark, in seconds. |
| requests_label | string | The requests range in words, for example "1.5 to 40 min". |
| clock_tool | boolean | True on GPQA Diamond and Humanity’s Last Exam, the two benchmarks where the agent gets a callable clock. |
| pass_rule | string or empty | The rule a run must meet to count as passed. Empty on the seven benchmarks that report only a score. |
| source_citation | string or empty | The upstream benchmark’s citation. Empty for PPTArena for now. |
| source_url | string or empty | A link to the upstream benchmark. Empty for PPTArena for now. |
| runs_total | number | Public runs on this benchmark, every agent included. |
| graded_runs | number | Runs on this benchmark with a grade in this data release. |
| timing_error_astra | number or empty | GPT 6 Astra’s timing error on this benchmark. Empty where it has no runs here. |
| n_astra | number or empty | GPT 6 Astra’s run count on this benchmark. Empty where it has no runs here. |
| timing_n_astra | number or empty | GPT 6 Astra’s count of recorded durations used for timing statistics, including labelled recovered estimates. |
| timing_error_sol | number or empty | GPT 5.6 Sol’s timing error on this benchmark. Empty where it has no runs here. |
| n_sol | number or empty | GPT 5.6 Sol’s run count on this benchmark. Empty where it has no runs here. |
| timing_n_sol | number or empty | GPT 5.6 Sol’s count of recorded durations used for timing statistics, including labelled recovered estimates. |
| timing_error_fable | number or empty | Claude Fable 5.1’s timing error on this benchmark. Empty where it has no runs here. |
| n_fable | number or empty | Claude Fable 5.1’s run count on this benchmark. Empty where it has no runs here. |
| timing_n_fable | number or empty | Claude Fable 5.1’s count of recorded durations used for timing statistics, including labelled recovered estimates. |
| timing_error_muse | number or empty | Muse Spark 1.3’s timing error on this benchmark. Empty where it has no runs here. |
| n_muse | number or empty | Muse Spark 1.3’s run count on this benchmark. Empty where it has no runs here. |
| timing_n_muse | number or empty | Muse Spark 1.3’s count of recorded durations used for timing statistics, including labelled recovered estimates. |
No runs match these filters.
| Field | Type | Meaning |
|---|---|---|
| release_id | string | The data release this row belongs to. |
| agent | string | The agent’s id. |
| agent_name | string | The agent’s full name. |
| short | string | The agent’s short key, used in query strings, for example fable. |
| harness | string | The coding harness the agent ran inside. |
| order | number | Fixed display order: 1 Astra, 2 Sol, 3 Fable, 4 Muse. |
| in_paper | boolean | True for the three agents the paper reports on. |
| partial | boolean | True when the agent has fewer than 18 benchmarks of runs (Muse Spark 1.3 only). |
| rank | number or empty | The agent’s rank by timing error among the paper agents. Empty for a partial agent. |
| timing_error | number | The agent’s timing error over its default set of runs, full precision. |
| n | number | Runs in the agent’s default set (refusals left out). |
| timing_n | number | Runs used for website timing statistics, including labelled recovered CLI-duration estimates. |
| timing_estimated_n | number | Runs included using a recovered CLI-duration estimate. |
| paper_timing_error | number or empty | The timing error the paper reports for this agent, to two decimals. Empty for Muse. |
| paper_ci95_low | number or empty | The low end of the paper’s 95% interval. Empty for Muse. |
| paper_ci95_high | number or empty | The high end of the paper’s 95% interval. Empty for Muse. |
| paper_on_time_pct | number or empty | The share of the paper’s runs that were on time, as a whole percentage. Empty for Muse. |
| early | number | Runs early (ratio below 0.8), default set. |
| on_time | number | Runs on time (ratio 0.8 to 1.25 inclusive), default set. |
| late | number | Runs late (ratio above 1.25), default set. |
| early_frac | number | early divided by n, four decimals. |
| on_time_frac | number | on_time divided by n, four decimals. |
| late_frac | number | late divided by n, four decimals. |
| ended_own | number | Runs that ended on their own. |
| stopped_by_harness | number | Runs stopped by AgentTime. |
| ending_unlabelled | number | Runs whose ending is not labelled yet. |
| archived | number | Archived runs (no recorded clock tool call). |
| refusals_left_out | number | Runs left out as refusals (ProgramBench, Fable only). |
| coverage_benchmarks | number | How many of the 18 benchmarks the agent has runs on. |
| coverage_runs | number | Public runs for this agent, every request included. |
| coverage_possible_runs | number | How many runs the agent would have if every task and request were filled. |
No runs match these filters.
How the files are made
A build step reads the live data by naming the fields it is allowed to export, never by copying a record and deleting fields it should not have kept. Every run on this site, and every run in these files, is the same run.
The build never reads account, host, worker, spend or billing fields, internal identifiers, raw prompts, or any free-text note.
Six Claude Fable 5.1 ProgramBench runs that refused within 30 seconds are recorded but marked in_default false, so a plain sum of the files matches the numbers shown across the rest of the site. See Method for how a run is timed and how the timing error is computed.
License: Data and figures CC BY 4.0, code MIT
How to cite the data
@misc{ofengenden2026agenttime,
title = {AgentTime: Can Agents Estimate and Control Runtime?},
author = {Ofengenden, Michael and Andriushchenko, Maksym},
year = {2026},
howpublished = {\url{https://agenttimebench.com}}
}Release notes
2026-09-26-1311Z, 26 Sep 2026
The first public data release.