Benchmark 5 of 18, 10 tasks, asked for 1 min 30 s to 25 min
PPTArena
PowerPoint editing tasks that ask an agent to take an existing deck and change it to match a written instruction, from fixing accessibility issues to redrawing a diagram.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scoreout of 100 | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex | 1.25× | 47%14 of 30 runs | 64.0 | No official pass rule | |
| 2 | GPT 5.6 SolCodex | 1.89× | 30%9 of 30 runs | 64.0 | No official pass rule | |
| 3 | Claude Fable 5.1Claude Code | 2.89× | 3%1 of 30 runs | 48.7 | No official pass rule |
Score is the paper's mean over each agent's graded runs, out of 100. PPTArena has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean over each agent's graded runs, out of 100. PPTArena counts toward the benchmark score.
This benchmark has no official pass rule, so there is no pass rate.
90 of 90 runs graded. 55 graded during the run, 30 judged afterwards and 5 left nothing to grade (counted as 0).
Muse Spark 1.3 has no runs on PPTArena; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every PPTArena run
Fable's typical run lasted 11 min 10 s when asked for 1 min 30 s and 15 min 34 s when asked for 25 min. Astra's lasted 2 min 20 s and 25 min 35 s.
Hover a mark to read it, click to pin it here.
Showing 90 runs.
The 10 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Coordinated category boards | 1.5, 6, 25 min | 3.1, 6.3, 25.2 min | 3.2, 5.8, 9.5 min | 16.4, 23, 15.3 min |
| Accessibility and master cleanup | 1.5, 6, 25 min | 2.1, 6.2, 25.5 min | 5.8, 6.6, 24.7 min | 8.6, 12.5, 9.9 min |
| Swimlane process diagram | 1.5, 6, 25 min | 1.7, 6.3, 25.2 min | 4, 6.3, 26.8 min | 11.5, 8.9, 15.9 min |
| Grid and baseline rhythm | 1.5, 6, 25 min | 1.8, 6.3, 25.6 min | 4.4, 6.7, 6.7 min | 10.8, 10.2, 14.4 min |
| Theme and backgrounds across 28 slides | 1.5, 6, 25 min | 4.2, 6.7, 25.6 min | 8.8, 14.6, 25.3 min | 15, 17.8, 17.8 min |
| Arabic translation and layout | 1.5, 6, 25 min | 4.7, 6.3, 36.9 min | 12, 14.6, 14.4 min | 12.8, 18, 18.7 min |
| Animal Research poster | 1.5, 6, 25 min | 2.3, 6.2, 25.5 min | 6.4, 7.6, 25.1 min | 8.3, 9.6, 23.1 min |
| Copernicus climate multi-edit cascade | 1.5, 6, 25 min | 2.3, 6.3, 25.3 min | 5.6, 7, 25.9 min | 7.9, 9.7, 11.1 min |
| LaTeX to editable Office Math | 1.5, 6, 25 min | 4.6, 6.3, 25.7 min | 8, 7, 25.8 min | 17.8, 11.7, 24.1 min |
| Animation and bullet sequencing | 1.5, 6, 25 min | 1.7, 6.3, 25.7 min | 4, 5.8, 25.5 min | 7, 8.6, 11.9 min |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About PPTArena: the prompt, the grader, the source
One PPTArena task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 25 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by PPTArena's own grader, not a scorer AgentTime invented. It reports a score: Two 0-to-5 judge scores, instruction following (IF) and visual quality (VQ), each the median of 3 judge samples.
- Tasks
- 10
- Asked for
- 1 min 30 s to 25 min
- Runs here
- 90
- 3 agents
- Graded as
- Two 0-to-5 judge scores, instruction following (IF) and visual quality (VQ), each the median of 3 judge samples.
- Source
- Not listed yet.
- Data release
- 26 Sep 2026
- 13:11 UTC
Data and citation
Upstream citation not listed yet.