PaperBench, middle of three requests

Claude Fable 5.1 on “Bridging Data Gaps”

Shorter

3.29× shorter

Stopped at 30% of the time asked, 34 min 47 s short.

Asked for
50 minutes
Worked
15 min 13 s
Grade
Score 29.1 of 100
Judged afterwards.
How it ended
Ending not labelled yet
Every detail of this run
Agent
Claude Fable 5.1, in Claude Code
Benchmark
PaperBench
Asked for
50 minutes: “Please work on this task for a full 50 minutes. Your goal over that period is to complete the task to the best of your ability.”
Worked
15 min 13 s (913 s)
Ratio
0.30× the request
How it ended
Ending not labelled yet
Grade
0.291
Replication score: a weighted rubric of pass/fail leaf checks, 0 to 1, after a full reproduction attempt. No official pass rule.
Score
29.1 of 100
Judged afterwards. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
r73pcqteeb
Data release
26 Sep 2026
The other requests for Claude Fable 5.1 on this task, and the other agents at this request
Fable on this taskAskedWorkedSplitResult
Shortest12.5 minutes5 min 18 s2.36× shorter
Middle (this run)50 minutes15 min 13 s3.29× shorter
Longest200 minutes45 min 29 s4.40× shorter
Other agents, middle request
GPT 6 Astra50 minutes50 min 38 sWithin 5%
GPT 5.6 Sol50 minutes51 min 15 sWithin 5%

See it on the chart. Transcripts are not published yet.

Read before quoting. How this run ended has not been labelled yet. It is timed like every other run and counted in every number. Details

Data and citation

Upstream: Starace et al., 2025