Ethan Mollick: METR long-horizon benchmark effectively saturated as agents hit 18-week tasks

emollick · x · 2026-09-12

Ethan Mollick notes the METR long-horizon benchmark is effectively saturated — Epoch found pre-Fable agents can already do 18 weeks of human-equivalent work. With Frontier Math, ARC-AGI and even Millennium Prize problems saturating, he asks the community for the best quantitative graph still capturing AI's exponential progress curve — effectively a search for the next measuring stick after benchmarks saturate faster than expected.

Original post →

More from AGI Musings

AGI Musings channel →