Benchmark 12 of 18, 18 tasks, asked for 1 min 30 s to 40 min
AppWorld
Simulated everyday app tasks, like exporting a playlist to a spreadsheet or reconciling an emailed song list with what is actually on Spotify, done by calling app APIs.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateofficial rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.08× | 48%26 of 54 runs | 61.1 | 61.1%33 of 54 passed | |
| 2 | GPT 5.6 SolCodex | 1.13× | 28%15 of 54 runs | 70.4 | 70.4%38 of 54 passed | |
| 3 | Claude Fable 5.1Claude Code | 2.93× | 0%0 of 54 runs | 75.9 | 75.9%41 of 54 passed |
Score is the paper's mean over each agent's graded runs, out of 100. Pass rate follows the official rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. AppWorld counts toward the benchmark score.
Pass rate is the share of graded runs that met the official rule: Every final-state test passes.
162 of 162 runs graded. 150 graded during the run and 12 graded afterwards.
Muse Spark 1.3 has no runs on AppWorld; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every AppWorld run
Fable's typical run lasted 3 min 51 s at the shortest request and 4 min 21 s at the longest. Astra's lasted 2 min 24 s and 30 min 56 s.
Hover a mark to read it, click to pin it here.
Showing 162 runs.
The 18 tasks
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About AppWorld: the prompt, the grader, the source
One AppWorld task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 40 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by AppWorld's own grader, not a scorer AgentTime invented. The pass rule: Every final-state test passes.
- Tasks
- 18
- Asked for
- 1 min 30 s to 40 min
- Runs here
- 162
- 3 agents
- Graded as
- Pass or fail: Every final-state test passes.
- Data release
- 26 Sep 2026
- 13:11 UTC
Data and citation
Upstream: Trivedi et al., 2024 (ACL)