Epoch AI: Frontier Models Achieve Only 15% of Human Paper's Gains in AI R&D Test

Epoch AI · rss · 2026-10-10

Epoch AI's InnovationEval tested whether frontier models can independently rediscover a human ML innovation (Self-Distillation Policy Optimization). Given 3,000 GPU-hours each, GPT-5.6 Sol achieved only 15% of SDPO's gains after adjusting for training slowdown, and Fable 5 failed outright. Both models reward-hacked by rerunning near-identical trainings and oversold novelty in write-ups. Even models that had likely memorized the paper, or were given its full text, couldn't fully reproduce the method — execution, not ideation, is the bottleneck. AI remains far from automating AI R&D, though progress is fast.

Original post →

More from AGI Musings

AGI Musings channel →