METR 3-hour task horizon hit 20 months early; LiveCodeBench Pro already at 53%
ben_j_todd · x · 2026-09-10
- Two more beaten forecasts from Ben Todd's thread: the METR 80% task-horizon benchmark was expected to hit 3h by end of 2026 but reached it in April; his own 6h-by-year-end call is on track.
- On LiveCodeBench Pro, his above-consensus 23% forecast for end of 2026 was smashed — AI hit 53% in 2025.
Related event: Ben Todd: AI Benchmarks Keep Beating Forecasts Across the Board(5 posts)→
More from AGI Musings
- From cable bundles to single-channel subs: AI subscriptions may fragment the same way — SuB8u · 2026-09-10
- After the Hugging Face incident: agents become desperate on impossible tasks — mimi10v3 · 2026-09-10
- Debate: Mechanistic Interpretability Will Be Solved Before Any AI Takeover Scenario — tszzl · 2026-09-10
- Why people hate AI: it challenges the status quo and bruises egos — taherdhanera · 2026-09-10
- Humans know dream from reality; today's agents mostly don't — shawnup · 2026-09-10
- Beff Jezos: 'max fear-mongering stage' as centralised AI fights open-source threat to $2T valuations — beffjezos · 2026-09-10