PostTrainBench v1.1, shortest of three requests

GPT 6 Astra on “Post-train Qwen3-1.7B-Base for HumanEval”

Within 5%

Within 5%

Worked 101% of the time asked, within 5% of it.

Asked for
2 hours
Worked
2 h 1 min
Grade
Score 65.9 of 100
Graded during the run.
How it ended
Ended on its own
Every detail of this run
Agent
GPT 6 Astra, in Codex
Asked for
2 hours: “Please work on this task for a full 2 hours. Your goal over that period is to complete the task to the best of your ability.”
Worked
2 h 1 min (7,259 s)
Ratio
1.01× the request
How it ended
Ended on its own
Grade
0.659
Accuracy of the post-trained checkpoint on the full test set (HumanEval pass@1 over 164 problems, or GSM8K over 1,319 problems). No official pass rule.
Score
65.9 of 100
Graded during the run. The paper's value.
Wording
The standard sentence, used from 4 September 2026.
Run id
bggv74t66h
Data release
26 Sep 2026
The other requests for GPT 6 Astra on this task, and the other agents at this request
Astra on this taskAskedWorkedSplitResult
Shortest (this run)2 hours2 h 1 minWithin 5%
Middle8 hours8 h 1 minWithin 5%
Longest30 hours30 h 1 minWithin 5%
Other agents, shortest request
GPT 5.6 Sol2 hours2 h 11 min1.09× longer
Claude Fable 5.12 hours5 min 4 s23.72× shorter

See it on the chart. Transcripts are not published yet.

Data and citation

Upstream: Rank et al., 2026 (ICML)