Prime Intellect Evaluates Autonomous AI Research Capabilities Across 18 Frontier Models
mariofilhoml · x · 2026-08-28
Prime Intellect released results from an experiment measuring autonomous AI research capabilities. They ran 153 autonomous runs across 18 frontier models using the nanoGPT optimizer speedrun, with runs lasting up to eight days. Results show a significant performance gap between models in experiment selection, execution, and interpretation. While no fundamentally new methods were generated, models like Claude Fable 5 and Opus 5 dramatically outperformed others.
More from Models
- Glitch Exposes Gemini's Internal Thoughts — Regular_Preference64 · 2026-08-28
- GLM-5.2 Introduces Monitors to Combat Reward Hacking in RL — burny_tech · 2026-08-28
- Anthropic Luna Max test shows generous limits, high speed — timpera · 2026-08-28
- User calls Grokbot 'terrible', cites missing tasks — krishnan · 2026-08-28
- Optimize Models to Think Less, Not Just Generate More Reasoning Tokens — abacaj · 2026-08-28
- ChatGPT keeps appending mysterious code to user chats — TheMoonMidas · 2026-08-28