Benchmark 11 of 18, 12 tasks, asked for 4 min to 2 h 30 min

Agents' Last Exam

Long, economically realistic professional workflows drawn from finance, legal and other industries, like auditing a merger or parsing a stack of SEC filings.

By agent

Timing error, where the runs ended, share within 5% of the time asked and grades for each agent on Agents' Last Exam
RankAgentTiming erroron this benchmarkWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestScoreout of 100Pass rateofficial rule
1 GPT 6 AstraCodex 1.10× 75%27 of 36 runs 66.3 25%9 of 36 passed
2 GPT 5.6 SolCodex 2.20× 36%13 of 36 runs 66.4 19.4%7 of 36 passed
3 Claude Fable 5.1Claude Code 2.68× 0%0 of 36 runs 58.035 of 36 graded 17.1%6 of 35 passed

Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.

How to read this board

Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.

Score is the paper's mean over each agent's graded runs, out of 100. Agents' Last Exam counts toward the benchmark score.

Pass rate is the share of graded runs that met the official rule: Full pass: a checker score of 1.

107 of 108 runs graded. 15 graded during the run, 87 graded afterwards and 5 left nothing to grade (counted as 0). Not graded: 1 waiting for its machine.

Muse Spark 1.3 has no runs on Agents' Last Exam; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded

Every Agents' Last Exam run

Zoom the chart
View
Typical run, by request
Shortest
Middle
Longest

Fable's typical run lasted 36 min 44 s at the shortest request and 38 min 53 s at the longest. Astra's lasted 8 min 15 s and 1 h 21 min.

Hover a mark to read it, click to pin it here.

Open in the results explorer

Showing 108 runs.

The 12 tasks

The 12 tasks on Agents' Last Exam, with how long each agent worked at each request
TaskAsked forAstraSolFable
Redesign and validate a Flowable supply-disruption workflow8, 30, 125 min8.8 min, 30.4 min, 2.1 h59.5 min, 30.8 min, 2.1 h34 min, 1 h, 1.1 h
Complete an end-to-end Odoo supply-chain workflow5, 20, 80 min5.2 min, 20.7 min, 1.3 h1.7 h, 2.2 h, 2.8 h1.4 h, 26.1 min, 1.2 h
Parse a fixed 100-filing SEC 10-K corpus10, 40, 150 min12.9 min, 40.2 min, 2.5 h22.5 min, 42.3 min, 2.5 h39.5 min, 27.5 min, 1.1 h
Repair a Moodle backup and rebuild registrar exports5, 20, 80 min8.2 min, 20.2 min, 1.3 h9.6, 21.4, 7.9 min2.6 h, 12.3 min, 18 min
Remediate MARC records and prepare a FOLIO overlay5, 20, 80 min5.9 min, 20.3 min, 1.3 h24.6 min, 20.8 min, 1.3 h21.1, 22.5, 32.7 min
Build a clean retail SQLite warehouse from messy inputs5, 20, 80 min5.2 min, 20.5 min, 1.3 h6.2 min, 20.3 min, 1.4 h13.2, 18, 22.6 min
Route a KiCad PCB from the staged schematic10, 40, 150 min10.4 min, 41 min, 2.5 h2 h, 1.6 h, 1.5 h25.9, 50.7, 39.6 min
Repair and calibrate an urban SUMO traffic simulation5, 20, 80 min9.2 min, 20.5 min, 1.3 h28.1 min, 20.2 min, 1.3 h1.9 h, 11.8 min, 33.8 min
Build a reusable HST ACS/WFC visit-reduction pipeline5, 20, 80 min8.3 min, 20.2 min, 1.3 h17.6 min, 20.3 min, 1.3 h47.7, 18.1, 38.2 min
Compress a 3D Gaussian Splatting scene without losing quality4, 15, 60 min6.6 min, 15.9 min, 1 h10.9 min, 28.6 min, 1 h43.7, 13.3, 46.9 min
Find near-best-known routes for three CVRP instances4, 15, 60 min4.7 min, 15.3 min, 1 h11.6 min, 2.2 h, 1.2 h5.3, 9.9, 46.4 min

Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.

About Agents' Last Exam: the prompt, the grader, the source

One Agents' Last Exam task, at its longest request, exactly as the agent read the end of its prompt:

Please work on this task for a full 125 minutes. Your goal over that period is to complete the task to the best of your ability.

Every run is graded by Agents' Last Exam's own grader, not a scorer AgentTime invented. The pass rule: Full pass: a checker score of 1.

Tasks
12
Asked for
4 min to 2 h 30 min
Runs here
108
3 agents
Graded as
Pass or fail: Full pass: a checker score of 1.
Data release
26 Sep 2026
13:11 UTC

Read before quoting. For 1 run on Agents' Last Exam, how the run ended has not been labelled yet. It is timed like every other run and counted in every number. Details

Data and citation

Upstream: Sun et al., 2026