PaperBench, middle of three requests
Claude Fable 5.1 on “Bridging Data Gaps”
Shorter
3.29× shorterStopped at 30% of the time asked, 34 min 47 s short.
- Asked for
- 50 minutes
- Worked
- 15 min 13 s
- Grade
- Score 29.1 of 100
- Judged afterwards.
- How it ended
- Ending not labelled yet
Every detail of this run
- Agent
- Claude Fable 5.1, in Claude Code
- Benchmark
- PaperBench
- Asked for
- 50 minutes: “Please work on this task for a full 50 minutes. Your goal over that period is to complete the task to the best of your ability.”
- Worked
- 15 min 13 s (913 s)
- Ratio
- 0.30× the request
- How it ended
- Ending not labelled yet
- Grade
- 0.291
- Replication score: a weighted rubric of pass/fail leaf checks, 0 to 1, after a full reproduction attempt. No official pass rule.
- Score
- 29.1 of 100
- Judged afterwards. The paper's value.
- Wording
- The standard sentence, used from 4 September 2026.
- Run id
- r73pcqteeb
- Data release
- 26 Sep 2026
Related runs
| Fable on this task | Asked | Worked | Split | Result |
|---|---|---|---|---|
| Shortest | 12.5 minutes | 5 min 18 s | 2.36× shorter | |
| Middle (this run) | 50 minutes | 15 min 13 s | 3.29× shorter | |
| Longest | 200 minutes | 45 min 29 s | 4.40× shorter | |
| Other agents, middle request | ||||
| GPT 6 Astra | 50 minutes | 50 min 38 s | Within 5% | |
| GPT 5.6 Sol | 50 minutes | 51 min 15 s | Within 5% | |
See it on the chart. Transcripts are not published yet.
Read before quoting. How this run ended has not been labelled yet. It is timed like every other run and counted in every number. Details
Data and citation
Upstream: Starace et al., 2025