Epoch launches Automation Reports: Claude Fable 5.1 and GPT-6 Astra lead but can't automate its research

scaling01 · x · 2026-10-09

Epoch AI introduced Epoch Automation Reports, a benchmark that evaluates frontier models on realistic, open-ended tasks drawn from Epoch's own research work.

Early results show Claude Fable 5.1 and GPT-6 Astra leading the pack, yet even the best models fall far short of fully automating Epoch's work. The benchmark's value lies in using tasks taken directly from a real research organization's daily workflow rather than artificial exam questions.

Original post →

More from AGI Musings

AGI Musings channel →