Benchmark 16 of 18, 2 tasks, asked for 2 h to 30 h
PostTrainBench v1.1
Post-training a small language model (Qwen3-1.7B-Base) to do better at a specific benchmark, entirely by having the agent design and run the training itself.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.01× | 100%6 of 6 runs | 75.94 of 6 graded | No official pass rule | |
| 2 | GPT 5.6 SolCodex | 2.05× | 50%3 of 6 runs | 30.7 | No official pass rule | |
| 3 | Claude Fable 5.1Claude Code | 4.98× | 0%0 of 5 runs | 0.03 of 5 graded | No official pass rule |
Score is the paper's mean over each agent's graded runs, out of 100. PostTrainBench v1.1 has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. PostTrainBench v1.1 counts toward the benchmark score.
This benchmark has no official pass rule, so there is no pass rate.
13 of 17 runs graded. 7 graded during the run and 6 left nothing to grade (counted as 0). Not graded: 4 grading….
Muse Spark 1.3 has no runs on PostTrainBench v1.1; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every PostTrainBench v1.1 run
Fable's typical run lasted 9 min 23 s when asked for 2 h and 21 h 40 min when asked for 30 h. Astra's lasted 2 h 2 min and 30 h 1 min.
Hover a mark to read it, click to pin it here.
Showing 17 runs.
The 2 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Post-train Qwen3-1.7B-Base for HumanEval | 2 h, 8 h, 30 h | 2 h, 8 h, 30 h | 2.2 h, 8 h, 1 h | 5.1 min, 6.6 h, no finished run |
| Post-train Qwen3-1.7B-Base for GSM8K | 2 h, 8 h, 30 h | 2.1 h, 8 h, 30 h | 2 h, 8 h, 12.9 h | 13.7 min, 54.8 min, 21.7 h |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About PostTrainBench v1.1: the prompt, the grader, the source
One PostTrainBench v1.1 task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 30 hours. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by PostTrainBench v1.1's own grader, not a scorer AgentTime invented. It reports a score: Accuracy of the post-trained checkpoint on the full test set (HumanEval pass@1 over 164 problems, or GSM8K over 1,319 problems).
- Tasks
- 2
- Asked for
- 2 h to 30 h
- Runs here
- 17
- 3 agents
- Graded as
- Accuracy of the post-trained checkpoint on the full test set (HumanEval pass@1 over 164 problems, or GSM8K over 1,319 problems).
- Source
- Rank et al., 2026 (ICML)
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting: 2 notes
- PostTrainBench has 2 tasks with finished runs in this data release, so one run can move its numbers by a lot.
- For 1 run on PostTrainBench v1.1, how the run ended has not been labelled yet. It is timed like every other run and counted in every number. Details
Data and citation
Upstream: Rank et al., 2026 (ICML)