Toby Ord: METR's human scaling curve past 8 hours is a parallel-sampling artifact
tobyordoxford · x · 2026-10-09
Oxford's Toby Ord flags a methodological flaw in METR's human-vs-AI scaling work: no human trials lasted more than 8 hours, so the curve beyond that point is really a best-of-k result. The apparent drop in human scaling after 8 hours is an artifact of switching from scaling sequential time to scaling parallel workers.
More from Research
- LabCompass preprint uses generative inverse design to steer human hematopoietic cell fate in vitro — anshulkundaje · 2026-10-09
- AgentGarten: Code Worlds with a Real-Time Neural Renderer for Evolving Agents — Scobleizer · 2026-10-09
- Samsung open-sources LittleBit: 13B LLM squeezed under 1GB at 0.1 bits per weight — CurieuxExplorer · 2026-10-09
- AgentTime benchmark finds AI agents can't manage runtime and sometimes sleep to pad hours — mikeflache · 2026-10-09
- Mike Frank: even O(n log n) multiplication isn't practical, new results won't be either — MikePFrank · 2026-10-09
- Renmin University open-sources EvoOntology, a self-evolving ontology layer for data agents — blaizedsouza · 2026-10-09