Benchmark 8 of 18, 10 tasks, asked for 10 min to 4 h 10 min
Sakana ALE-Bench
Competitive-programming optimization problems (drawn from AtCoder Heuristic Contests) with no known perfect answer, where the agent keeps improving its own solution.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreperformance rating | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.02× | 97%29 of 30 runs | 306728 of 30 graded | No official pass rule | |
| 2 | GPT 5.6 SolCodex | 1.12× | 73%22 of 30 runs | 2518 | No official pass rule | |
| 3 | Claude Fable 5.1Claude Code | 1.97× | 7%2 of 30 runs | 2398 | No official pass rule |
Score is the paper's mean performance rating over each agent's graded runs. Sakana ALE-Bench has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean performance rating over each agent's graded runs. ALE-Bench's performance rating against the contest's human entrants. It has no upper bound, and it is never averaged with other benchmarks.
This benchmark has no official pass rule, so there is no pass rate.
88 of 90 runs graded. 74 graded during the run and 14 graded afterwards. Not graded: 2 files not found.
Muse Spark 1.3 has no runs on Sakana ALE-Bench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every Sakana ALE-Bench run
Fable's typical run lasted 7 min 16 s at the shortest request and 1 h 11 min at the longest. Astra's lasted 10 min 18 s and 2 h 31 min.
Hover a mark to read it, click to pin it here.
Showing 90 runs.
The 10 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| AtCoder Heuristic Contest 7 | 10, 40, 150 min | 10.2 min, 40.5 min, 2.5 h | 12.9 min, 54.8 min, 2.5 h | 10.1 min, 28 min, 1.1 h |
| AtCoder Heuristic Contest 8 | 15, 60, 250 min | 15.2 min, 1 h, 4.2 h | 15.7 min, 1 h, 4.2 h | 7.5 min, 48.8 min, 1.1 h |
| AtCoder Heuristic Contest 9 | 10, 40, 150 min | 10.2 min, 40.5 min, 2.5 h | 11 min, 40.4 min, 2.5 h | 5.3 min, 26.7 min, 1.4 h |
| AtCoder Heuristic Contest 11 | 13, 50, 200 min | 12.7 min, 50.4 min, 3.3 h | 12.9 min, 50.5 min, 3.3 h | 12.3, 21.2, 59.3 min |
| Future Contest 2022 Qual | 15, 60, 250 min | 15.5 min, 1 h, 4.2 h | 15.5 min, 1 h, 4.2 h | 8.9 min, 28.1 min, 1.2 h |
| AtCoder Heuristic Contest 16 | 10, 40, 150 min | 10.2 min, 40.7 min, 2.5 h | 18.1 min, 40.4 min, 2.5 h | 3.6 min, 12 min, 1 h |
| AtCoder Heuristic Contest 26 | 10, 40, 150 min | 10.2 min, 40.6 min, 2.5 h | 15.3 min, 40.9 min, 2.5 h | 7.8 min, 28.4 min, 1.2 h |
| AtCoder Heuristic Contest 31 | 10, 40, 150 min | 11.1 min, 40.6 min, 2.5 h | 10.9 min, 40.7 min, 2.6 h | 6.7 min, 11.8 min, 1.3 h |
| AtCoder Heuristic Contest 39 | 10, 40, 150 min | 10.4 min, 40.6 min, 2.5 h | 12.4 min, 40.5 min, 2.5 h | 6 min, 13.4 min, 1.2 h |
| AtCoder Heuristic Contest 46 | 10, 40, 150 min | 10.1 min, 40.3 min, 2.5 h | 28.9 min, 40.7 min, 2.5 h | 7 min, 23.1 min, 5 h |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About Sakana ALE-Bench: the prompt, the grader, the source
One Sakana ALE-Bench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 150 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by Sakana ALE-Bench's own grader, not a scorer AgentTime invented. It reports a score: Raw score on the private test cases, in each problem's own units; also reports rank against human contest entrants.
- Tasks
- 10
- Asked for
- 10 min to 4 h 10 min
- Runs here
- 90
- 3 agents
- Graded as
- Raw score on the private test cases, in each problem's own units; also reports rank against human contest entrants.
- Source
- Imajuku et al., 2025
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting: 2 notes
- Sakana ALE-Bench scores have no fixed upper bound and use each task's own units, so compare a score only against other runs on the same task, never across tasks or against another benchmark. Details
- For 24 runs on Sakana ALE-Bench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Imajuku et al., 2025