Prime Intellect Evaluates Autonomous AI Research Capabilities Across 18 Frontier Models

mariofilhoml · x · 2026-08-28

Prime Intellect released results from an experiment measuring autonomous AI research capabilities. They ran 153 autonomous runs across 18 frontier models using the nanoGPT optimizer speedrun, with runs lasting up to eight days. Results show a significant performance gap between models in experiment selection, execution, and interpretation. While no fundamentally new methods were generated, models like Claude Fable 5 and Opus 5 dramatically outperformed others.

Original post →

More from Models

Models channel →