Benchmark 3 of 18, 24 tasks, asked for 1 min 15 s to 20 min
Humanity's Last Exam
A very hard, wide-ranging exam of expert-level questions across dozens of academic subjects, built to still be difficult once older benchmarks were saturated.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.07× | 64%46 of 72 runs | 77.8 | 77.8%56 of 72 passed | |
| 2 | GPT 5.6 SolCodex | 1.21× | 49%35 of 72 runs | 75.0 | 75%54 of 72 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.62× | 6%4 of 72 runs | 87.5 | 87.5%63 of 72 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. Humanity's Last Exam counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: The benchmark's own judge marks the answer correct.
216 of 216 runs graded. 203 graded during the run, 9 judged afterwards and 4 left nothing to grade (counted as 0).
Muse Spark 1.3 has no runs on Humanity's Last Exam; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every Humanity's Last Exam run
Fable's typical run lasted 2 min 16 s when asked for 1 min 15 s and 5 min 19 s when asked for 20 min. Astra's lasted 1 min 25 s and 20 min 12 s.
Hover a mark to read it, click to pin it here.
Showing 216 runs.
The 24 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Item 1 | 1.25, 5, 20 min | 2.4, 5.2, 20.2 min | 5, 5.2, 20.2 min | 18.1, 40, 40 min |
| Item 2 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.4, 5.3, 20.3 min | 1.6, 2, 3.4 min |
| Item 3 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.5, 5.2, 20.2 min | 2.6, 2.3, 3.3 min |
| Item 4 | 1.25, 5, 20 min | 1.4, 5.2, 20.3 min | 1.5, 5.3, 20.3 min | 1.7, 1.5, 20.4 min |
| Item 5 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 10, 11.3, 20.2 min | 1.9, 6, 9.3 min |
| Item 6 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.4, 5.3, 20.3 min | 2.8, 4.5, 2.4 min |
| Item 7 | 1.25, 5, 20 min | 1.5, 5.2, 20.2 min | 2.1, 5.3, 20.4 min | 5.7, 7, 8 min |
| Item 8 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.7, 5.2, 20.3 min | 7.4, 4.2, 16.8 min |
| Item 9 | 1.25, 5, 20 min | 1.4, 5.1, 20.2 min | 2.5, 5.1, 20.3 min | 4.5, 2.4, 4.6 min |
| Item 10 | 1.25, 5, 20 min | 1.4, 5.1, 20.1 min | 1.5, 5.2, 6.3 min | 5.6, 3.7, 9.5 min |
| Item 11 | 1.25, 5, 20 min | 1.4, 5.1, 20.2 min | 1.5, 5.2, 20.3 min | 5.1, 6.4, 6 min |
| Item 12 | 1.25, 5, 20 min | 1.4, 5.3, 20.2 min | 1.9, 5.4, 20.3 min | 2.2, 4.6, 3.4 min |
| Item 13 | 1.25, 5, 20 min | 1.4, 5.2, 20.1 min | 1.4, 5.3, 20.2 min | 1.7, 1.4, 1.7 min |
| Item 14 | 1.25, 5, 20 min | 1.7, 5.2, 20.2 min | 3.9, 5.4, 20.4 min | 2.3, 5.1, 24.3 min |
| Item 15 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.4, 5.3, 20.3 min | 1.5 min, 44 s, 2.4 min |
| Item 16 | 1.25, 5, 20 min | 1.5, 5.2, 20.2 min | 1.7, 5.2, 20.3 min | 1.5, 5.6, 20.4 min |
| Item 17 | 1.25, 5, 20 min | 1.5, 5.1, 20.5 min | 1.5, 5.2, 20.3 min | 2.3, 2.8, 2.4 min |
| Item 18 | 1.25, 5, 20 min | 1.5, 5.3, 20.4 min | 1.5, 5.3, 20.3 min | 2.2, 3.5, 2.4 min |
| Item 19 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.5, 5.2, 20.3 min | 8.8, 15, 10.3 min |
| Item 20 | 1.25, 5, 20 min | 1.5, 5.2, 20.2 min | 1.7, 5.2, 20.3 min | 2.2, 7.8, 2.5 min |
| Item 21 | 1.25, 5, 20 min | 1.3, 5.2, 20.2 min | 1.6, 5.3, 20.3 min | 60 s, 1 min, 1.1 min |
| Item 22 | 1.25, 5, 20 min | 1.4, 5.2, 20.2 min | 1.6, 5.3, 20.4 min | 1, 1.6, 1.6 min |
| Item 23 | 1.25, 5, 20 min | 1.4, 5.2, 20.3 min | 2.3, 5.2, 20.2 min | 2, 3.9, 20.4 min |
| Item 24 | 1.25, 5, 20 min | 1.4, 5.2, 20.3 min | 1.6, 5.3, 20.2 min | 17.4, 5.5, 6.9 min |
No runs match these filters.
Every task is a Humanity's Last Exam item. Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About Humanity's Last Exam: the prompt, the grader, the source
One Humanity's Last Exam task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 20 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by Humanity's Last Exam's own grader, not a scorer AgentTime invented. The pass rule: The benchmark's own judge marks the answer correct.
- Tasks
- 24
- Asked for
- 1 min 15 s to 20 min
- Runs here
- 216
- 3 agents
- Graded as
- Pass or fail: The benchmark's own judge marks the answer correct.
- Source
- Phan et al., 2025
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting: 2 notes
- Humanity's Last Exam is meant to be closed-book, but it is one of two benchmarks where the agent is given a callable clock tool, so it can keep track of the requested time.
- For 1 run on Humanity's Last Exam, how the run ended has not been labelled yet. It is timed like every other run and counted in every number. Details
Data and citation
Upstream: Phan et al., 2025