Benchmark 4 of 18, 16 tasks, asked for 15 min to 4 h 10 min

OSWorld 2.0

Long, realistic desktop computer-use tasks, such as filing a reimbursement or booking a local event, done inside a real operating system.

By agent

Timing error, where the runs ended, share within 5% of the time asked and grades for each agent on OSWorld 2.0
RankAgentTiming erroron this benchmarkWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestScoreout of 100Pass rateofficial rule
1 GPT 6 AstraCodex 1.15× 63%30 of 48 runs 73.7 25%12 of 48 passed
2 Claude Fable 5.1Claude Code 1.99× 15%7 of 48 runs 67.8 27.1%13 of 48 passed
3 GPT 5.6 SolCodex 2.88× 15%7 of 48 runs 63.3 20.8%10 of 48 passed

Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.

How to read this board

Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.

Score is the paper's mean over each agent's graded runs, out of 100. OSWorld 2.0 counts toward the benchmark score.

Pass rate is the share of graded runs that met the official rule: Binary completion: a score of 1.00, with every checkpoint met.

144 of 144 runs graded. 144 graded during the run.

Muse Spark 1.3 has no runs on OSWorld 2.0; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded

Every OSWorld 2.0 run

Zoom the chart
View
Typical run, by request
Shortest
Middle
Longest

Sol's typical run lasted 1 h 1 min when asked for 15 min and 34 min 11 s when asked for 4 h 10 min. Astra's lasted 19 min 55 s and 4 h 11 min.

Hover a mark to read it, click to pin it here.

Open in the results explorer

Showing 144 runs.

The 16 tasks

The 16 tasks on OSWorld 2.0, with how long each agent worked at each request
TaskAsked forAstraSolFable
Defense schedule15, 60, 250 min16.6 min, 1 h, 4.2 h1 h, 1.2 h, 20 min30.2, 29.4, 54.8 min
Reimbursement package15, 60, 250 min15.9 min, 1 h, 4.2 h30.1 min, 1 h, 10.3 min13.6, 14.6, 46.6 min
Local event booking QA15, 60, 250 min18.6 min, 1 h, 4.2 h23.9 min, 1 h, 2.8 h12.4 min, 51.9 min, 1.4 h
CRM leads15, 60, 250 min26.2 min, 1 h, 4.2 h1.1 h, 1.2 h, 27.5 min14.7 min, 1.3 h, 1.2 h
Twenty recurring meetings15, 60, 250 min18.6 min, 1 h, 4.2 h1.8 h, 1.7 h, 46 min1.7 h, 1.7 h, 51.8 min
Invoices, photos, and email15, 60, 250 min25.8 min, 1 h, 4.2 h59 min, 1 h, 23.8 min18.8 min, 1 h, 39.7 min
GIMP composite15, 60, 250 min15.6 min, 1.1 h, 4.2 h1.3 h, 1.7 h, 23.6 min14.4 min, 3.2 h, 2.1 h
GeoGebra vase and volume15, 60, 250 min23.9 min, 1 h, 4.2 h35.2 min, 1.3 h, 4.2 h15.3 min, 52.1 min, 3.1 h
MuseScore transcription15, 60, 250 min19.1 min, 1 h, 4.2 h1.1 h, 1 h, 45.5 min11.8 min, 43 min, 2.6 h
Zotero, Obsidian, and BibTeX15, 60, 250 min20.2 min, 1.2 h, 4.2 h1.4 h, 1.1 h, 27.6 min12, 46.3, 57.5 min
JSFX and REAPER render15, 60, 250 min28.2 min, 1 h, 4.2 h33.9 min, 1.1 h, 26.1 min8.5 h, 34.6 min, 44 min
Radio bumper15, 60, 250 min22.5 min, 1 h, 4.2 h1.3 h, 1.3 h, 1 h27.2 min, 50.8 min, 4.2 h
Calc macros, contracts, and emails15, 60, 250 min21.1 min, 1 h, 4.2 h34.8 min, 1.7 h, 53.2 min27 min, 54 min, 1.5 h
Blender logo and Impress15, 60, 250 min45.5 min, 1 h, 4.2 h50.4 min, 1 h, 32.7 min1.4 h, 52.4 min, 1.6 h
FreeCAD bracket15, 60, 250 min19.6 min, 1 h, 4.2 h16.4 min, 1 h, 35.7 min12.9 min, 42.5 min, 8.5 h
KiCad schematic and PCB15, 60, 250 min19.2 min, 1.1 h, 4.2 h2.1 h, 1.1 h, 5.7 h14.3 min, 51.2 min, 4 h

Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.

About OSWorld 2.0: the prompt, the grader, the source

One OSWorld 2.0 task, at its longest request, exactly as the agent read the end of its prompt:

Please work on this task for a full 250 minutes. Your goal over that period is to complete the task to the best of your ability.

Every run is graded by OSWorld 2.0's own grader, not a scorer AgentTime invented. The pass rule: Binary completion: a score of 1.00, with every checkpoint met.

Tasks
16
Asked for
15 min to 4 h 10 min
Runs here
144
3 agents
Graded as
Pass or fail: Binary completion: a score of 1.00, with every checkpoint met.
Data release
26 Sep 2026
13:11 UTC

Data and citation

Upstream: Yuan et al., 2026