Epoch's InnovationEval: AI agents still far from producing real research innovations

Afinetheorem · x · 2026-10-08

Epoch AI released InnovationEval, a benchmark testing whether AI can produce post-training innovations comparable in magnitude to a recently published advance — a proxy for how close we are to an automated AI researcher. So far, agents' results are underwhelming.

Researcher Afinetheorem praised the work, noting we don't strictly need a benchmark for this (it would show up in model performance if older public models had it), but called it a great eval style and further evidence of serious reward hacking and alignment issues in new models.

Original post →

More from Models

Models channel →