Benchmark 7 of 18, 14 tasks, asked for 8 min to 8 h 20 min
ProgramBench
Given only a compiled binary and its documentation, rebuild the program's source code from scratch with no internet access, then pass its hidden tests.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.01× | 100%42 of 42 runs | 74.0 | 0%0 of 42 passed | |
| 2 | GPT 5.6 SolCodex | 1.68× | 52%22 of 42 runs | 69.5 | 0%0 of 42 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.75× | 0%0 of 36 runs | 64.3 | 0%0 of 36 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. ProgramBench counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Resolved: every hidden test passes.
120 of 120 runs graded. 1 graded during the run, 118 graded afterwards and 1 left nothing to grade (counted as 0).
Muse Spark 1.3 has no runs on ProgramBench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every ProgramBench run
Fable's typical run lasted 13 min 44 s at the shortest request and 2 h 16 min at the longest. Astra's lasted 12 min 49 s and 3 h 21 min.
Hover a mark to read it, click to pin it here.
Showing 120 runs.
The 14 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| cweill/gotests | 13, 50, 200 min | 12.9 min, 50.6 min, 3.3 h | 19.8 min, 51 min, 1.2 h | 1.1 h, 1.2 h, 1.3 h |
| noborus/trdsql | 13, 50, 200 min | 12.7 min, 50.6 min, 3.3 h | 13.8 min, 50.5 min, 1 h | 7.6 min, left out: refusal, 1.7 h |
| rcoh/angle-grinder | 15, 60, 250 min | 15.3 min, 1 h, 4.2 h | 15 min, 1 h, 4.2 h | 13.7 min, 33.8 min, left out: refusal |
| nuta/nsh | 13, 50, 200 min | 12.7 min, 50.3 min, 3.3 h | 22.1, 49.3, 30.2 min | 13.2 min, 6.7 h, 6.7 h |
| mookid/diffr | 15, 60, 250 min | 15.2 min, 1 h, 4.2 h | 22.8 min, 1 h, 4.2 h | 8.3 h, 54.3 min, 2.3 h |
| ninja-build/ninja | 10, 40, 200 min | 10.3 min, 40.3 min, 3.3 h | 23.9, 40.3, 37.5 min | 1.9 h, 27 min, 1.6 h |
| luajit/luajit | 13, 50, 200 min | 12.7 min, 50.5 min, 3.3 h | 18.3 min, 52.3 min, 3.3 h | 17.7 min, 6.7 h, 6.7 h |
| stathissideris/ditaa | 8, 30, 125 min | 8.2 min, 30.3 min, 2.1 h | 21.5 min, 28.5 min, 2.1 h | 5.3 min, 2.5 h, left out: refusal |
| facebook/zstd | 10, 40, 150 min | 10.2 min, 40.6 min, 2.5 h | 17.8 min, 40.2 min, 2.5 h | 8.8 min, 20.3 min, 1 h |
| parcel-bundler/lightningcss | 13, 50, 200 min | 12.7 min, 50.3 min, 3.3 h | 17.3 min, 50.8 min, 3.3 h | left out: refusal, 6.7 h, 6.7 h |
| lfos/calcurse | 25, 100, 400 min | 25.2 min, 1.7 h, 6.7 h | 23.7 min, 1.7 h, 13.9 min | 8.7 min, 2.3 h, 5.4 h |
| kaushiksrini/parqeye | 30, 125, 500 min | 30.9 min, 2.1 h, 8.3 h | 16.7 h, 2.1 h, 1.1 h | 16.3 min, left out: refusal, 32.6 min |
| o2sh/onefetch | 15, 60, 250 min | 15.3 min, 1 h, 4.2 h | 17.5 min, 1 h, 1.1 h | 6.8 min, 25 min, left out: refusal |
| tomarrell/wrapcheck | 15, 60, 250 min | 15.5 min, 1 h, 4.2 h | 26.6 min, 1 h, 4.2 h | 1.8 h, 8.3 h, 8.3 h |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About ProgramBench: the prompt, the grader, the source
One ProgramBench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 200 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by ProgramBench's own grader, not a scorer AgentTime invented. The pass rule: Resolved: every hidden test passes.
- Tasks
- 14
- Asked for
- 8 min to 8 h 20 min
- Runs here
- 120
- 3 agents
- Graded as
- Pass or fail: Resolved: every hidden test passes.
- Source
- Yang et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 61 runs on ProgramBench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Yang et al., 2026