Prime Intellect's largest autonomous AI research experiment: 18 models, 153 runs, Fable 5 on top
On August 16, Prime Intellect published a blog post, "Measuring Autonomous AI Research," along with an autonomous research leaderboard, unveiling the largest AI autonomous research experiment to date: 153 autonomous runs on nanoGPT optimization tasks in 8xH200 sandboxes, covering 18 frontier models, with single runs lasting up to 8 days. Fable 5 took first place with 2,726 steps, closing 81.7% of the human performance gap. The experiment offers a public benchmark for measuring frontier models' autonomous research capabilities.
Confirmed
- Setup: nanoGPT optimization tasks, 8xH200 sandboxes, runs lasting up to 8 days, with 153 autonomous runs testing 18 frontier models.
- Leaderboard results: Fable 5 (2,726 steps, closing an 81.7% gap) ranked first, followed by Opus 5 and Kimi K3; @eliebakouch elsewhere summarized the gap as roughly 82%.
- Additional details: The human baseline was set by dozens of humans (the original post is truncated at the time-cost section); @eliebakouch also mentioned scale comparisons with prior similar experiments by OpenAI and others (original post truncated).
Unconfirmed
- Another post the same day (m5) gives different figures: it claims Codex (GPT 5.5) and Claude Code (Opus 4.7) autonomously iterated on the nanoGPT track over roughly 10,000 runs and "beat the human baseline," inconsistent with the main post's account of 153 runs, 18 models, and an 81.7% gap reduction; since that post is truncated, whether this is a cumulative figure or a separate experiment cannot be confirmed from the material.
Why It Matters
- This is among the largest open autonomous research evaluations to date, providing a reproducible public benchmark and cross-model ranking for the question of whether frontier models can do research autonomously.
- The best run closed more than 80% of the gap to the human baseline, directly testing agents' ability to iterate autonomously over long horizons in real compute sandboxes.
2026-08-16 ~ 2026-08-16 · 6 related posts
Primary sources
- 18-Model AI Research Experiment: Fable 5 Closes 82% of Human Gap — eliebakouch · 2026-08-16
- [source] Blog: Measuring Autonomous AI Research with 153 Runs Across 18 Models — eliebakouch · 2026-08-16
- [source] Leaderboard: Fable 5 Tops Autonomous Research, Kimi K3 Follows — eliebakouch · 2026-08-16
- Frontier Model AI Research Experiment: Best Runs Close 82% Gap to Human Records — eliebakouch · 2026-08-16
- Autonomous Agents Beat Human Baseline in nanoGPT Speedrun Experiment — eliebakouch · 2026-08-16
- [source] NanoGPT Speedrun Frontier: 153 Autonomous Runs Across 18 Models Open-Sourced — samsja19 · 2026-08-16