Ethan Mollick: METR long-horizon benchmark effectively saturated as agents hit 18-week tasks
emollick · x · 2026-09-12
Ethan Mollick notes the METR long-horizon benchmark is effectively saturated — Epoch found pre-Fable agents can already do 18 weeks of human-equivalent work. With Frontier Math, ARC-AGI and even Millennium Prize problems saturating, he asks the community for the best quantitative graph still capturing AI's exponential progress curve — effectively a search for the next measuring stick after benchmarks saturate faster than expected.
More from AGI Musings
- Gary Marcus pushes back on prediction that Anthropic defectors will return within 6 months — GaryMarcus · 2026-09-12
- OpenAI Wasn't 'Out of Control' — It Was a Calculated Trade-off, Argues User — kchonyc · 2026-09-12
- EA grant platform hires FTX's Caroline Ellison, drawing fire over alignment judgment — zetalyrae · 2026-09-12
- RL debate: REINFORCE is both policy gradient descent and a synthetic data method — jessi_cata · 2026-09-12
- OpenAI was never 'out of control' — it could always shut down its data centers, argues critic — kchonyc · 2026-09-12
- ARC-AGI cost collapse: DeepSeek-V4-Flash hits 87% at $0.021/task vs o3's $4,500 — inductionheads · 2026-09-12