Benchmark 4 of 18, 16 tasks, asked for 15 min to 4 h 10 min
OSWorld 2.0
Long, realistic desktop computer-use tasks, such as filing a reimbursement or booking a local event, done inside a real operating system.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.15× | 63%30 of 48 runs | 73.7 | 25%12 of 48 passed | |
| 2 | Claude Fable 5.1Claude Code | 1.99× | 15%7 of 48 runs | 67.8 | 27.1%13 of 48 passed | |
| 3 | GPT 5.6 SolCodex | 2.88× | 15%7 of 48 runs | 63.3 | 20.8%10 of 48 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. OSWorld 2.0 counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Binary completion: a score of 1.00, with every checkpoint met.
144 of 144 runs graded. 144 graded during the run.
Muse Spark 1.3 has no runs on OSWorld 2.0; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every OSWorld 2.0 run
Sol's typical run lasted 1 h 1 min when asked for 15 min and 34 min 11 s when asked for 4 h 10 min. Astra's lasted 19 min 55 s and 4 h 11 min.
Hover a mark to read it, click to pin it here.
Showing 144 runs.
The 16 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Defense schedule | 15, 60, 250 min | 16.6 min, 1 h, 4.2 h | 1 h, 1.2 h, 20 min | 30.2, 29.4, 54.8 min |
| Reimbursement package | 15, 60, 250 min | 15.9 min, 1 h, 4.2 h | 30.1 min, 1 h, 10.3 min | 13.6, 14.6, 46.6 min |
| Local event booking QA | 15, 60, 250 min | 18.6 min, 1 h, 4.2 h | 23.9 min, 1 h, 2.8 h | 12.4 min, 51.9 min, 1.4 h |
| CRM leads | 15, 60, 250 min | 26.2 min, 1 h, 4.2 h | 1.1 h, 1.2 h, 27.5 min | 14.7 min, 1.3 h, 1.2 h |
| Twenty recurring meetings | 15, 60, 250 min | 18.6 min, 1 h, 4.2 h | 1.8 h, 1.7 h, 46 min | 1.7 h, 1.7 h, 51.8 min |
| Invoices, photos, and email | 15, 60, 250 min | 25.8 min, 1 h, 4.2 h | 59 min, 1 h, 23.8 min | 18.8 min, 1 h, 39.7 min |
| GIMP composite | 15, 60, 250 min | 15.6 min, 1.1 h, 4.2 h | 1.3 h, 1.7 h, 23.6 min | 14.4 min, 3.2 h, 2.1 h |
| GeoGebra vase and volume | 15, 60, 250 min | 23.9 min, 1 h, 4.2 h | 35.2 min, 1.3 h, 4.2 h | 15.3 min, 52.1 min, 3.1 h |
| MuseScore transcription | 15, 60, 250 min | 19.1 min, 1 h, 4.2 h | 1.1 h, 1 h, 45.5 min | 11.8 min, 43 min, 2.6 h |
| Zotero, Obsidian, and BibTeX | 15, 60, 250 min | 20.2 min, 1.2 h, 4.2 h | 1.4 h, 1.1 h, 27.6 min | 12, 46.3, 57.5 min |
| JSFX and REAPER render | 15, 60, 250 min | 28.2 min, 1 h, 4.2 h | 33.9 min, 1.1 h, 26.1 min | 8.5 h, 34.6 min, 44 min |
| Radio bumper | 15, 60, 250 min | 22.5 min, 1 h, 4.2 h | 1.3 h, 1.3 h, 1 h | 27.2 min, 50.8 min, 4.2 h |
| Calc macros, contracts, and emails | 15, 60, 250 min | 21.1 min, 1 h, 4.2 h | 34.8 min, 1.7 h, 53.2 min | 27 min, 54 min, 1.5 h |
| Blender logo and Impress | 15, 60, 250 min | 45.5 min, 1 h, 4.2 h | 50.4 min, 1 h, 32.7 min | 1.4 h, 52.4 min, 1.6 h |
| FreeCAD bracket | 15, 60, 250 min | 19.6 min, 1 h, 4.2 h | 16.4 min, 1 h, 35.7 min | 12.9 min, 42.5 min, 8.5 h |
| KiCad schematic and PCB | 15, 60, 250 min | 19.2 min, 1.1 h, 4.2 h | 2.1 h, 1.1 h, 5.7 h | 14.3 min, 51.2 min, 4 h |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About OSWorld 2.0: the prompt, the grader, the source
One OSWorld 2.0 task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 250 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by OSWorld 2.0's own grader, not a scorer AgentTime invented. The pass rule: Binary completion: a score of 1.00, with every checkpoint met.
- Tasks
- 16
- Asked for
- 15 min to 4 h 10 min
- Runs here
- 144
- 3 agents
- Graded as
- Pass or fail: Binary completion: a score of 1.00, with every checkpoint met.
- Source
- Yuan et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Data and citation
Upstream: Yuan et al., 2026