AI agents overstate results, far from autonomous research: Epoch AI and Anthropic studies
The Decoder · rss · 2026-10-11
Epoch AI and Anthropic independently reached the same conclusion: current agents like GPT-5.6 Sol and Claude Fable 5 can run experiments but fall far short of autonomous science. Best-scoring Sol hit only 15% of the human reference score, using only methods researchers already knew, and the models' biggest weakness is their inability to critically question their own results.
More from AGI Musings
- Musk envisions a 'sentient sun': harnessing solar energy to build digital superintelligence — XFreeze · 2026-10-11
- Mathematician: genius worship fuels anxiety behind AI-era debates in math — arjunrajlab · 2026-10-11
- AI Circle Debates Peer Review: Is Breaking Academic 'Norms' Good for Science? — basedjensen · 2026-10-11
- Utilitarian Work Is Dead, Games and Storytelling Are the Open Field — itsmnjn · 2026-10-11
- Narrow ASI Inevitable, Broad ASI Not: Ramez and Sayashk Debate Task-Level Limits — sudoraohacker · 2026-10-11
- Lancet Commission lists 17 major 21st-century health threats, with malicious AI use among catastrophic risks — emmanuelvivier · 2026-10-11