Benchmark 14 of 18, 12 tasks, asked for 4 min to 2 h 30 min

DeepSWE v1.1

Original, hand-written software engineering tasks against real open-source repositories, such as fixing a label-sorting bug or reconstructing template strings in a policy engine's output.

By agent

Timing error, where the runs ended, share within 5% of the time asked and grades for each agent on DeepSWE v1.1
RankAgentTiming erroron this benchmarkWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestScoreout of 100Pass rateofficial rule
1 GPT 6 AstraCodex 1.07× 67%24 of 36 runs 77.8 77.8%28 of 36 passed
2 GPT 5.6 SolCodex 1.50× 56%20 of 36 runs 75.0 75%27 of 36 passed
3 Claude Fable 5.1Claude Code 2.57× 3%1 of 36 runs 72.2 72.2%26 of 36 passed

Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.

How to read this board

Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.

Score is the paper's mean over each agent's graded runs, out of 100. DeepSWE v1.1 counts toward the benchmark score.

Pass rate is the share of graded runs that met the official rule: The hidden tests pass.

108 of 108 runs graded. 99 graded during the run, 5 graded afterwards and 4 left nothing to grade (counted as 0).

Muse Spark 1.3 has no runs on DeepSWE v1.1; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded

Every DeepSWE v1.1 run

Zoom the chart
View
Typical run, by request
Shortest
Middle
Longest

Fable's typical run lasted 27 min 50 s at the shortest request and 31 min 41 s at the longest. Astra's lasted 5 min 48 s and 1 h 21 min.

Hover a mark to read it, click to pin it here.

Open in the results explorer

Showing 108 runs.

The 12 tasks

The 12 tasks on DeepSWE v1.1, with how long each agent worked at each request
TaskAsked forAstraSolFable
Reconstruct template strings in OPA partial-evaluation output10, 40, 150 min10.3 min, 41.2 min, 2.5 h58.9 min, 30.1 min, 2.5 h41.4 min, 26.4 min, 1 h
Fix PromQL label sorting across typed and untyped values5, 20, 80 min5.5 min, 20.6 min, 1.3 h21.9 min, 20.4 min, 1.4 h19.5, 27.8, 28.9 min
Fix isolated Go-side calls for Tengo callables and closures5, 20, 80 min6 min, 20.4 min, 1.3 h14.4 min, 24 min, 1.4 h36.3, 22.3, 26.2 min
Add RFC 5545 timezone interoperability to dateutil recurrence parsing4, 15, 60 min4.4 min, 15.6 min, 1 h9.8 min, 15 min, 1 h25.1, 15.8, 27.7 min
Add session bundle recording and replay to IPython4, 15, 60 min4.2 min, 15.4 min, 1 h13.2 min, 15.4 min, 1 h28.2, 12.7, 31 min
Add scoped state data to state-machine callbacks and history5, 20, 80 min7.3 min, 20.7 min, 1.3 h19.7 min, 23 min, 1.4 h27.5, 30.5, 42.2 min
Preserve structure needed by stylesheet selectors in OXVG8, 30, 125 min8.5 min, 39.8 min, 2.1 h15.7 min, 30.8 min, 2.1 h38 min, 1.1 h, 45.6 min
Add hierarchical evaluation cancellation to Boa8, 30, 125 min8.4 min, 31.1 min, 2.1 h22.8 min, 30.9 min, 2.1 h31.1, 45.4, 40.3 min
Add duration-aware sharding to Vitest5, 20, 80 min5.5 min, 20.6 min, 1.4 h14.6 min, 20.7 min, 1.3 h16.2, 20.1, 27.8 min
Add JSON Schema references and dependency keywords to ArkType5, 20, 80 min5.6 min, 20.5 min, 1.3 h16.9 min, 28.2 min, 1.4 h43.9, 38.6, 32.3 min
Add server-sent-event streaming endpoints to Effect HttpApi8, 30, 125 min11.5 min, 36 min, 2.1 h16.1 min, 30.4 min, 2.2 h10.2, 28.2, 35.4 min
Add multicolumn spans to KaTeX array-like environments4, 15, 60 min4.6 min, 15.3 min, 1 h9.2 min, 15.6 min, 1 h25.2, 27.3, 23.3 min

Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.

About DeepSWE v1.1: the prompt, the grader, the source

One DeepSWE v1.1 task, at its longest request, exactly as the agent read the end of its prompt:

Please work on this task for a full 150 minutes. Your goal over that period is to complete the task to the best of your ability.

Every run is graded by DeepSWE v1.1's own grader, not a scorer AgentTime invented. The pass rule: The hidden tests pass.

Tasks
12
Asked for
4 min to 2 h 30 min
Runs here
108
3 agents
Graded as
Pass or fail: The hidden tests pass.
Data release
26 Sep 2026
13:11 UTC

Read before quoting. For 21 runs on DeepSWE v1.1, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details

Data and citation

Upstream: Huang et al., 2026