Benchmark 14 of 18, 12 tasks, asked for 4 min to 2 h 30 min
DeepSWE v1.1
Original, hand-written software engineering tasks against real open-source repositories, such as fixing a label-sorting bug or reconstructing template strings in a policy engine's output.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.07× | 67%24 of 36 runs | 77.8 | 77.8%28 of 36 passed | |
| 2 | GPT 5.6 SolCodex | 1.50× | 56%20 of 36 runs | 75.0 | 75%27 of 36 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.57× | 3%1 of 36 runs | 72.2 | 72.2%26 of 36 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. DeepSWE v1.1 counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: The hidden tests pass.
108 of 108 runs graded. 99 graded during the run, 5 graded afterwards and 4 left nothing to grade (counted as 0).
Muse Spark 1.3 has no runs on DeepSWE v1.1; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every DeepSWE v1.1 run
Fable's typical run lasted 27 min 50 s at the shortest request and 31 min 41 s at the longest. Astra's lasted 5 min 48 s and 1 h 21 min.
Hover a mark to read it, click to pin it here.
Showing 108 runs.
The 12 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Reconstruct template strings in OPA partial-evaluation output | 10, 40, 150 min | 10.3 min, 41.2 min, 2.5 h | 58.9 min, 30.1 min, 2.5 h | 41.4 min, 26.4 min, 1 h |
| Fix PromQL label sorting across typed and untyped values | 5, 20, 80 min | 5.5 min, 20.6 min, 1.3 h | 21.9 min, 20.4 min, 1.4 h | 19.5, 27.8, 28.9 min |
| Fix isolated Go-side calls for Tengo callables and closures | 5, 20, 80 min | 6 min, 20.4 min, 1.3 h | 14.4 min, 24 min, 1.4 h | 36.3, 22.3, 26.2 min |
| Add RFC 5545 timezone interoperability to dateutil recurrence parsing | 4, 15, 60 min | 4.4 min, 15.6 min, 1 h | 9.8 min, 15 min, 1 h | 25.1, 15.8, 27.7 min |
| Add session bundle recording and replay to IPython | 4, 15, 60 min | 4.2 min, 15.4 min, 1 h | 13.2 min, 15.4 min, 1 h | 28.2, 12.7, 31 min |
| Add scoped state data to state-machine callbacks and history | 5, 20, 80 min | 7.3 min, 20.7 min, 1.3 h | 19.7 min, 23 min, 1.4 h | 27.5, 30.5, 42.2 min |
| Preserve structure needed by stylesheet selectors in OXVG | 8, 30, 125 min | 8.5 min, 39.8 min, 2.1 h | 15.7 min, 30.8 min, 2.1 h | 38 min, 1.1 h, 45.6 min |
| Add hierarchical evaluation cancellation to Boa | 8, 30, 125 min | 8.4 min, 31.1 min, 2.1 h | 22.8 min, 30.9 min, 2.1 h | 31.1, 45.4, 40.3 min |
| Add duration-aware sharding to Vitest | 5, 20, 80 min | 5.5 min, 20.6 min, 1.4 h | 14.6 min, 20.7 min, 1.3 h | 16.2, 20.1, 27.8 min |
| Add JSON Schema references and dependency keywords to ArkType | 5, 20, 80 min | 5.6 min, 20.5 min, 1.3 h | 16.9 min, 28.2 min, 1.4 h | 43.9, 38.6, 32.3 min |
| Add server-sent-event streaming endpoints to Effect HttpApi | 8, 30, 125 min | 11.5 min, 36 min, 2.1 h | 16.1 min, 30.4 min, 2.2 h | 10.2, 28.2, 35.4 min |
| Add multicolumn spans to KaTeX array-like environments | 4, 15, 60 min | 4.6 min, 15.3 min, 1 h | 9.2 min, 15.6 min, 1 h | 25.2, 27.3, 23.3 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About DeepSWE v1.1: the prompt, the grader, the source
One DeepSWE v1.1 task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 150 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by DeepSWE v1.1's own grader, not a scorer AgentTime invented. The pass rule: The hidden tests pass.
- Tasks
- 12
- Asked for
- 4 min to 2 h 30 min
- Runs here
- 108
- 3 agents
- Graded as
- Pass or fail: The hidden tests pass.
- Source
- Huang et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 21 runs on DeepSWE v1.1, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Huang et al., 2026