OpenAI internal models show ~15-min 80% time-horizon on real research tasks, far below METR

ben_j_todd · x · 2026-09-29

Ben Todd weighs in on an internal OpenAI evaluation finding: the 80% time-horizon of its strongest internal models on real research tasks is only about 15 minutes — well below METR's generic software-engineering numbers or UK AISI's cyber-task results.

He offers two caveats: the results are only 8x behind METR, roughly a year of progress; and models often still succeed at 64h+ tasks, providing some evidence on creative research work. With a fairly smooth curve, the natural projection is continued across-the-board progress.

Original post →

More from AGI Musings

AGI Musings channel →