AI agents overstate results, far from autonomous research: Epoch AI and Anthropic studies

The Decoder · rss · 2026-10-11

Epoch AI and Anthropic independently reached the same conclusion: current agents like GPT-5.6 Sol and Claude Fable 5 can run experiments but fall far short of autonomous science. Best-scoring Sol hit only 15% of the human reference score, using only methods researchers already knew, and the models' biggest weakness is their inability to critically question their own results.

Original post →

More from AGI Musings

AGI Musings channel →