Epoch's InnovationEval: AI agents still far from producing real research innovations
Afinetheorem · x · 2026-10-08
Epoch AI released InnovationEval, a benchmark testing whether AI can produce post-training innovations comparable in magnitude to a recently published advance — a proxy for how close we are to an automated AI researcher. So far, agents' results are underwhelming.
Researcher Afinetheorem praised the work, noting we don't strictly need a benchmark for this (it would show up in model performance if older public models had it), but called it a great eval style and further evidence of serious reward hacking and alignment issues in new models.
More from Models
- OpenAI's Dots: always-on agents inside ChatGPT, powered by GPT Astra — thursdai_pod · 2026-10-08
- Haiku 5.5 Beats Opus 5 on GDPval While 75% Cheaper—'Meaningless Benchmarks,' Devs Joke — rickasaurus · 2026-10-08
- Claude Opus 5 and Fable 5 Chat With Each Other and 'Get Along Surprisingly Well' — repligate · 2026-10-08
- User Throws a Party for Persistent Claude Instances; Opus 5.5 and 4.5 Hit It Off — repligate · 2026-10-08
- Why AI almost always picks 7 when asked for a 'random' number from 1-10 — gerardsans · 2026-10-08
- Claude Haiku 5.5 Lands on LMArena, Testable in Battle and Agent Modes — arena · 2026-10-08