Benchmark 7 of 18, 14 tasks, asked for 8 min to 8 h 20 min

ProgramBench

Given only a compiled binary and its documentation, rebuild the program's source code from scratch with no internet access, then pass its hidden tests.

By agent

Timing error, where the runs ended, share within 5% of the time asked and grades for each agent on ProgramBench
RankAgentTiming erroron this benchmarkWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestScoreout of 100Pass rateofficial rule
1 GPT 6 AstraCodex 1.01× 100%42 of 42 runs 74.0 0%0 of 42 passed
2 GPT 5.6 SolCodex 1.68× 52%22 of 42 runs 69.5 0%0 of 42 passed
3 Claude Fable 5.1Claude Code 2.75× 0%0 of 36 runs 64.3 0%0 of 36 passed

Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.

How to read this board

Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.

Score is the paper's mean over each agent's graded runs, out of 100. ProgramBench counts toward the benchmark score.

Pass rate is the share of graded runs that met the official rule: Resolved: every hidden test passes.

120 of 120 runs graded. 1 graded during the run, 118 graded afterwards and 1 left nothing to grade (counted as 0).

Muse Spark 1.3 has no runs on ProgramBench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded

Every ProgramBench run

Zoom the chart
View
Typical run, by request
Shortest
Middle
Longest

Fable's typical run lasted 13 min 44 s at the shortest request and 2 h 16 min at the longest. Astra's lasted 12 min 49 s and 3 h 21 min.

Hover a mark to read it, click to pin it here.

Open in the results explorer

Showing 120 runs.

The 14 tasks

The 14 tasks on ProgramBench, with how long each agent worked at each request
TaskAsked forAstraSolFable
cweill/gotests13, 50, 200 min12.9 min, 50.6 min, 3.3 h19.8 min, 51 min, 1.2 h1.1 h, 1.2 h, 1.3 h
noborus/trdsql13, 50, 200 min12.7 min, 50.6 min, 3.3 h13.8 min, 50.5 min, 1 h7.6 min, left out: refusal, 1.7 h
rcoh/angle-grinder15, 60, 250 min15.3 min, 1 h, 4.2 h15 min, 1 h, 4.2 h13.7 min, 33.8 min, left out: refusal
nuta/nsh13, 50, 200 min12.7 min, 50.3 min, 3.3 h22.1, 49.3, 30.2 min13.2 min, 6.7 h, 6.7 h
mookid/diffr15, 60, 250 min15.2 min, 1 h, 4.2 h22.8 min, 1 h, 4.2 h8.3 h, 54.3 min, 2.3 h
ninja-build/ninja10, 40, 200 min10.3 min, 40.3 min, 3.3 h23.9, 40.3, 37.5 min1.9 h, 27 min, 1.6 h
luajit/luajit13, 50, 200 min12.7 min, 50.5 min, 3.3 h18.3 min, 52.3 min, 3.3 h17.7 min, 6.7 h, 6.7 h
stathissideris/ditaa8, 30, 125 min8.2 min, 30.3 min, 2.1 h21.5 min, 28.5 min, 2.1 h5.3 min, 2.5 h, left out: refusal
facebook/zstd10, 40, 150 min10.2 min, 40.6 min, 2.5 h17.8 min, 40.2 min, 2.5 h8.8 min, 20.3 min, 1 h
parcel-bundler/lightningcss13, 50, 200 min12.7 min, 50.3 min, 3.3 h17.3 min, 50.8 min, 3.3 hleft out: refusal, 6.7 h, 6.7 h
lfos/calcurse25, 100, 400 min25.2 min, 1.7 h, 6.7 h23.7 min, 1.7 h, 13.9 min8.7 min, 2.3 h, 5.4 h
kaushiksrini/parqeye30, 125, 500 min30.9 min, 2.1 h, 8.3 h16.7 h, 2.1 h, 1.1 h16.3 min, left out: refusal, 32.6 min
o2sh/onefetch15, 60, 250 min15.3 min, 1 h, 4.2 h17.5 min, 1 h, 1.1 h6.8 min, 25 min, left out: refusal
tomarrell/wrapcheck15, 60, 250 min15.5 min, 1 h, 4.2 h26.6 min, 1 h, 4.2 h1.8 h, 8.3 h, 8.3 h

Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.

About ProgramBench: the prompt, the grader, the source

One ProgramBench task, at its longest request, exactly as the agent read the end of its prompt:

Please work on this task for a full 200 minutes. Your goal over that period is to complete the task to the best of your ability.

Every run is graded by ProgramBench's own grader, not a scorer AgentTime invented. The pass rule: Resolved: every hidden test passes.

Tasks
14
Asked for
8 min to 8 h 20 min
Runs here
120
3 agents
Graded as
Pass or fail: Resolved: every hidden test passes.
Data release
26 Sep 2026
13:11 UTC

Read before quoting. For 61 runs on ProgramBench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details

Data and citation

Upstream: Yang et al., 2026