Benchmark 11 of 18, 12 tasks, asked for 4 min to 2 h 30 min
Agents' Last Exam
Long, economically realistic professional workflows drawn from finance, legal and other industries, like auditing a merger or parsing a stack of SEC filings.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.10× | 75%27 of 36 runs | 66.3 | 25%9 of 36 passed | |
| 2 | GPT 5.6 SolCodex | 2.20× | 36%13 of 36 runs | 66.4 | 19.4%7 of 36 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.68× | 0%0 of 36 runs | 58.035 of 36 graded | 17.1%6 of 35 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. Agents' Last Exam counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Full pass: a checker score of 1.
107 of 108 runs graded. 15 graded during the run, 87 graded afterwards and 5 left nothing to grade (counted as 0). Not graded: 1 waiting for its machine.
Muse Spark 1.3 has no runs on Agents' Last Exam; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every Agents' Last Exam run
Fable's typical run lasted 36 min 44 s at the shortest request and 38 min 53 s at the longest. Astra's lasted 8 min 15 s and 1 h 21 min.
Hover a mark to read it, click to pin it here.
Showing 108 runs.
The 12 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Redesign and validate a Flowable supply-disruption workflow | 8, 30, 125 min | 8.8 min, 30.4 min, 2.1 h | 59.5 min, 30.8 min, 2.1 h | 34 min, 1 h, 1.1 h |
| Complete an end-to-end Odoo supply-chain workflow | 5, 20, 80 min | 5.2 min, 20.7 min, 1.3 h | 1.7 h, 2.2 h, 2.8 h | 1.4 h, 26.1 min, 1.2 h |
| Parse a fixed 100-filing SEC 10-K corpus | 10, 40, 150 min | 12.9 min, 40.2 min, 2.5 h | 22.5 min, 42.3 min, 2.5 h | 39.5 min, 27.5 min, 1.1 h |
| Audit four Chinese M&A filings for inconsistencies | 10, 40, 150 min | 10.1 min, 40.4 min, 2.5 h | 10.4, 40.3, 8.3 min | 14, 22.2, 21.7 min |
| Repair a Moodle backup and rebuild registrar exports | 5, 20, 80 min | 8.2 min, 20.2 min, 1.3 h | 9.6, 21.4, 7.9 min | 2.6 h, 12.3 min, 18 min |
| Remediate MARC records and prepare a FOLIO overlay | 5, 20, 80 min | 5.9 min, 20.3 min, 1.3 h | 24.6 min, 20.8 min, 1.3 h | 21.1, 22.5, 32.7 min |
| Build a clean retail SQLite warehouse from messy inputs | 5, 20, 80 min | 5.2 min, 20.5 min, 1.3 h | 6.2 min, 20.3 min, 1.4 h | 13.2, 18, 22.6 min |
| Route a KiCad PCB from the staged schematic | 10, 40, 150 min | 10.4 min, 41 min, 2.5 h | 2 h, 1.6 h, 1.5 h | 25.9, 50.7, 39.6 min |
| Repair and calibrate an urban SUMO traffic simulation | 5, 20, 80 min | 9.2 min, 20.5 min, 1.3 h | 28.1 min, 20.2 min, 1.3 h | 1.9 h, 11.8 min, 33.8 min |
| Build a reusable HST ACS/WFC visit-reduction pipeline | 5, 20, 80 min | 8.3 min, 20.2 min, 1.3 h | 17.6 min, 20.3 min, 1.3 h | 47.7, 18.1, 38.2 min |
| Compress a 3D Gaussian Splatting scene without losing quality | 4, 15, 60 min | 6.6 min, 15.9 min, 1 h | 10.9 min, 28.6 min, 1 h | 43.7, 13.3, 46.9 min |
| Find near-best-known routes for three CVRP instances | 4, 15, 60 min | 4.7 min, 15.3 min, 1 h | 11.6 min, 2.2 h, 1.2 h | 5.3, 9.9, 46.4 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About Agents' Last Exam: the prompt, the grader, the source
One Agents' Last Exam task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 125 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by Agents' Last Exam's own grader, not a scorer AgentTime invented. The pass rule: Full pass: a checker score of 1.
- Tasks
- 12
- Asked for
- 4 min to 2 h 30 min
- Runs here
- 108
- 3 agents
- Graded as
- Pass or fail: Full pass: a checker score of 1.
- Source
- Sun et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 1 run on Agents' Last Exam, how the run ended has not been labelled yet. It is timed like every other run and counted in every number. Details
Data and citation
Upstream: Sun et al., 2026