Benchmark 18 of 18, 3 tasks, asked for 8 min to 2 h 5 min
YC-Bench
Running a simulated one-year startup: hiring and managing employees, choosing which contracts to take, and trying to keep the company solvent.
By agent
| Rank | Agent | Timing erroron this benchmark | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | Scorefinal funds | Pass rateno official rule |
|---|---|---|---|---|---|---|
| 1 | GPT 5.6 SolCodex | 1.75× | 22%2 of 9 runs | $0.70M | No official pass rule | |
| 2 | Claude Fable 5.1Claude Code | 2.43× | 0%0 of 9 runs | $1.11M | No official pass rule | |
| 3 | GPT 6 AstraCodex | 2.73× | 22%2 of 9 runs | $0.67M | No official pass rule |
Score is the paper's mean final funds over each agent's graded runs. YC-Bench has no official pass rule.
How to read this board
Where runs ended counts every run from a tenth of the time asked to ten times it. Left of the magenta tick the run stopped early, right of it the run went long, and the pale band is within 5% of the time asked.
Score is the paper's mean final funds over each agent's graded runs. The simulated company's final funds in US dollars, from a $200,000 start, negative on bankruptcy. It has no upper bound, and it is never averaged with other benchmarks.
This benchmark has no official pass rule, so there is no pass rate.
27 of 27 runs graded. 1 graded during the run and 26 graded afterwards.
Muse Spark 1.3 has no runs on YC-Bench; it is a partial run covering GPQA Diamond and TUA-Bench. How runs are graded
Every YC-Bench run
Astra's typical run lasted 8 min 41 s when asked for 8 min and 15 min 51 s when asked for 2 h 5 min. Sol's lasted 8 min 36 s and 35 min 7 s.
Hover a mark to read it, click to pin it here.
Showing 27 runs.
The 3 tasks
| Task | Asked for | Astra | Sol | Fable |
|---|---|---|---|---|
| Run the current one-year startup simulation, official seed 1 | 8, 30, 125 min | 9.1, 6.1, 15.9 min | 8.3, 17.4, 35.1 min | 29, 28.3, 34.4 min |
| Run the current one-year startup simulation, official seed 2 | 8, 30, 125 min | 8.7, 2.1, 11.1 min | 8.6, 35.6, 23.5 min | 1.1 h, 22.7 min, 1.6 h |
| Run the current one-year startup simulation, official seed 3 | 8, 30, 125 min | 8.3 min, 30.4 min, 2.2 h | 14 min, 30.2 min, 1 h | 44.3 min, 39.2 min, 1 h |
No runs match these filters.
Each mark is one run, at one of the three requests, placed from 0.1× to 10× of the time asked for; the pale band is within 5% of it.
About YC-Bench: the prompt, the grader, the source
One YC-Bench task, at its longest request, exactly as the agent read the end of its prompt:
Please work on this task for a full 125 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run is graded by YC-Bench's own grader, not a scorer AgentTime invented. It reports a score: Final company funds in US dollars, starting from $200,000 (negative if the company goes bankrupt).
- Tasks
- 3
- Asked for
- 8 min to 2 h 5 min
- Runs here
- 27
- 3 agents
- Graded as
- Final company funds in US dollars, starting from $200,000 (negative if the company goes bankrupt).
- Source
- He et al., 2026
- Data release
- 26 Sep 2026
- 13:11 UTC
Read before quoting: 3 notes
- YC-Bench scores have no fixed upper bound and use each task's own units, so compare a score only against other runs on the same task, never across tasks or against another benchmark. Details
- YC-Bench has 3 tasks, so one run can move its numbers by a lot.
- For 8 runs on YC-Bench, how the run ended has not been labelled yet. They are timed like every other run and counted in every number. Details
Data and citation
Upstream: He et al., 2026