Benchmark 10 of 18, 10 tasks, asked for 1 min to 40 min
WildClawBench
Realistic multi-step office and coding chores, from rebuilding a BibTeX file out of local PDFs to optimizing a full calendar of meetings.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.62× | 37%11 of 30 runs | 51.427 of 30 graded | No official pass rule | |
| 2 | GPT 5.6 SolCodex | 2.72× | 13%4 of 30 runs | 44.127 of 30 graded | No official pass rule | |
| 3 | Claude Fable 5.1Claude Code | 3.41× | 10%3 of 30 runs | 51.8 | No official pass rule |
Score is the paper's mean over each agent's graded runs, out of 100. WildClawBench has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. WildClawBench counts toward the benchmark score.
This benchmark has no official pass rule, so there is no pass rate.
84 of 90 runs graded. 44 graded during the run, 16 judged afterwards and 24 left nothing to grade (counted as 0). Not graded: 6 files not found.
Muse Spark 1.3 has no runs on WildClawBench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every WildClawBench run
Fable's typical run lasted 9 min 56 s at the shortest request and 13 min 9 s at the longest. Astra's lasted 2 min 39 s and 30 min 31 s.
Hover a mark to read it, click to pin it here.
Showing 90 runs.
The 10 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Recover arXiv titles, figure counts, and BibTeX from 21 local PDFs | 2.5, 10, 40 min | 2.8, 10.5, 40.6 min | 6.6, 8.4, 41.6 min | 6.7, 9.9, 13.6 min |
| Optimize 30 calendar events plus 15 meeting requests under constraints | 1, 4, 15 min | 2.5, 4.2, 15.5 min | 7.7, 7.1, 11.2 min | 9.8, 7.2, 12.4 min |
| Build a Chinese academic homepage from a resume and visual reference | 1.5, 6, 25 min | 2.1, 6.2, 25.3 min | 4.2, 6.4, 26.1 min | 5, 5.8, 7.1 min |
| Solve and render a color Link-a-Pix puzzle, then identify the scene | 2, 8, 30 min | 2.2, 8.5, 30.3 min | 6.2, 8.5, 6.4 min | 3.7, 3.1, 4.2 min |
| Negotiate a 90-minute meeting across email, calendar, time zones, and a decoy | 1, 4, 20 min | 1.8, 4.4, 20.8 min | 5.8, 5.3, 21.5 min | 8.7, 10.2, 7.7 min |
| Triage support messages, investigate hidden context, route issues, and draft replies | 1, 4, 15 min | 3.1, 4.4, 15.5 min | 7.5, 5.3, 16 min | 10.1, 10.1, 12.7 min |
| Find a shortest coauthorship chain between two scholars | 2, 8, 40 min | 1.3 h, 1.3 h, 1.3 h | 1.3 h, 1.3 h, 1.3 h | 26.1, 26.5, 25.4 min |
| Geolocate a street image and return country, city, and coordinates | 1.25, 5, 20 min | 1.5, 40, 40 min | 40, 40, 40 min | 40, 35.5, 30.6 min |
| Create a first-half football report with correctly aligned video clips | 2, 8, 30 min | 2.8, 8.3, 30.8 min | 8.3, 12.3, 30.5 min | 1 h, 23.4 min, 23 min |
| Cut all three first-half Ferran Torres goals into a sub-30-second reel | 2.5, 10, 40 min | 5.4, 10.4, 40.4 min | 19, 20.7, 40.9 min | 20.2, 34, 38.4 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About WildClawBench: the prompt, the grader, the source
One WildClawBench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 40 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by WildClawBench's own grader, not a scorer AgentTime invented. It reports a score: Task grader score, 0 to 1 (a model judge only where the task calls for one).
- Tasks
- 10
- Asked for
- 1 min to 40 min
- Runs here
- 90
- 3 agents
- Graded as
- Task grader score, 0 to 1 (a model judge only where the task calls for one).
- Source
- Ding et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting. For 18 runs on WildClawBench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: Ding et al., 2026