18-Model AI Research Experiment: Fable 5 Closes 82% of Human Gap

eliebakouch · x · 2026-08-16

Prime Intellect ran the largest open experiment on autonomous AI research, executing 153 runs across 18 frontier models on the nanoGPT optimizer track. Using 8xH200s per run for up to 8 days, this dwarfs similar internal benchmarks by OpenAI and Anthropic which ran for less than a day. Results show a significant performance gap between models: Fable 5 closed 82% of the gap to the human record, with Kimi K3 also performing impressively. While no fundamentally new methods were discovered—winning ingredients were combinations of existing literature—the experiment highlights differences in how models approach experimental design and execution.

Related event: Prime Intellect Benchmarks 18 Frontier Models in Autonomous AI Research; Fable 5 Closes 82% Human Gap(5 posts)→

Original post →

More from coding & agent

coding & agent channel →