Blog: Measuring Autonomous AI Research with 153 Runs Across 18 Models

eliebakouch · x · 2026-08-16

Prime Intellect published 'Measuring Autonomous AI Research' to evaluate how well frontier models can conduct research. They ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 models, with individual runs lasting up to eight days. The authors argue that while speedrun methods may not inherently scale to real training, the tight feedback loops make it an interesting testbed. The most striking finding is the significant gap between models, evident in their choice of experiments, execution, and interpretation. No run produced a fundamentally new method, yet Fable 5 and Opus 5 performed dramatically better than the rest.

Related event: Largest Autonomous AI Research Experiment Released(7 posts)→

Original post →

More from coding & agent

coding & agent channel →