OpenAI internal models show ~15-min 80% time-horizon on real research tasks, far below METR
ben_j_todd · x · 2026-09-29
Ben Todd weighs in on an internal OpenAI evaluation finding: the 80% time-horizon of its strongest internal models on real research tasks is only about 15 minutes — well below METR's generic software-engineering numbers or UK AISI's cyber-task results.
He offers two caveats: the results are only 8x behind METR, roughly a year of progress; and models often still succeed at 64h+ tasks, providing some evidence on creative research work. With a fairly smooth curve, the natural projection is continued across-the-board progress.
More from AGI Musings
- François Fleuret: living in the AI age will feel like being royalty, commanding smarter minds — francoisfleuret · 2026-09-29
- Are AI slowdown calls a cover for progress that's already stalling? — Comfortable_Key_1346 · 2026-09-29
- Meta launches Muse for Small Business, connecting the AI agent to Shopify, Slack and 15 more tools — gaganghotra_ · 2026-09-29
- "Superhuman general intelligence already exists — it's humans cooperating," Reddit argues — EC36339 · 2026-09-29
- Eric Elliott resurfaces Leanpub podcast on AI Driven Development, consciousness and economics — ericelliott_ · 2026-09-29
- Anthropic maps multiagent system risks; researcher likens it to sociology, not chemistry — mattturck · 2026-09-29