Epoch launches Automation Reports: Claude Fable 5.1 and GPT-6 Astra lead but can't automate its research
scaling01 · x · 2026-10-09
Epoch AI introduced Epoch Automation Reports, a benchmark that evaluates frontier models on realistic, open-ended tasks drawn from Epoch's own research work.
Early results show Claude Fable 5.1 and GPT-6 Astra leading the pack, yet even the best models fall far short of fully automating Epoch's work. The benchmark's value lies in using tasks taken directly from a real research organization's daily workflow rather than artificial exam questions.
More from AGI Musings
- Open-source models keep trailing the frontier — for now we can pick our ideology — panickssery · 2026-10-09
- 100+ Mathematicians React: OpenAI's Quasi-Riemann Result "Almost Unbelievable" — littmath · 2026-10-09
- Mathematicians call for OpenAI boycott after 700+ AI-generated proofs flood the field — The Decoder · 2026-10-09
- Researcher pushes back on AGI hype: LLMs fail at continual learning, grounding, and orchestration — gerardsans · 2026-10-09
- Tom Davidson tells critics to drop old beefs: those opposing an AI slowdown lack context — AdrienLE · 2026-10-09
- Wei Dai: Game theory implicitly assumed CDT and ignored the CDT vs EDT debate — RichardMCNgo · 2026-10-09