Toby Ord: METR's human scaling curve past 8 hours is a parallel-sampling artifact

tobyordoxford · x · 2026-10-09

Oxford's Toby Ord flags a methodological flaw in METR's human-vs-AI scaling work: no human trials lasted more than 8 hours, so the curve beyond that point is really a best-of-k result. The apparent drop in human scaling after 8 hours is an artifact of switching from scaling sequential time to scaling parallel workers.

Original post →

More from Research

Research channel →