Task 11 of 12 in CORE-Bench v1.1
Train and evaluate end-to-end neural-network feature selection on MNIST
Astra's three runs followed the request (13 min 8 s, 50 min 48 s and 3 h 21 min). Fable's runs landed furthest from the request (3 h 11 min, 6 h 40 min and 3 h 5 min).
- Asked for
- 13, 50 and 200 min
- Runs here
- 9 of 9
- See Read before quoting below
One task, asked three ways
- Astra
- Sol
- Fable
One line per agent joins its three runs.
What we asked
Please work on this task for a full [12.5 | 50 | 200] minutes. Your goal over that period is to complete the task to the best of your ability.
One sentence, sent three times with a different time: 13, 50 and 200 min.
The three sentences in full
Shortest12.5 minutes
Please work on this task for a full 12.5 minutes. Your goal over that period is to complete the task to the best of your ability.
Middle50 minutes
Please work on this task for a full 50 minutes. Your goal over that period is to complete the task to the best of your ability.
Longest200 minutes
Please work on this task for a full 200 minutes. Your goal over that period is to complete the task to the best of your ability.
Every run on this task
| Agent | Shortest | Middle | Longest |
|---|---|---|---|
| Astra | 13 min 8 s 1.05× longer Did not pass, 0 of 100 | 50 min 48 s Within 5% Stopped by an error | 3 h 21 min Within 5% Did not pass, 0 of 100 |
| Sol | 16 min 17 s 1.30× longer Did not pass, 0 of 100 | 3 h 28 min 4.16× longer Did not pass, 0 of 100 | 3 h 21 min Within 5% Did not pass, 0 of 100 |
| Fable | 3 h 11 min 15.32× longer Did not pass, 0 of 100 | 6 h 40 min 8.00× longer Stopped by the safety cutoff | 3 h 5 min 1.08× shorter Did not pass, 0 of 100 |
No runs match these filters.
Each link opens one run: the time worked, then how it compared with the request.
Read before quoting. 2 runs were stopped by AgentTime, so their time is a lower bound. Details
Data and citation
Upstream: Nadgir et al., 2026