o3's time horizon hits ~110 minutes yet no system clears 31% on DeepScholar-Bench
le_james94 · x · 2026-09-16
Noting that DeepSeekMath-V2 makes verification the product, this post highlights how measurement is splitting in two: METR puts o3's 50% time horizon near 110 minutes of human work, doubling roughly every 7 months, while on DeepScholar-Bench no tested system clears a 31% geometric mean at writing related-work sections. Both numbers are real.
Related event: Reading Stanford CS329A in Depth: Generators Have Outpaced Verifiers(5 posts)→
More from Research
- A new type of LM that outputs probabilities: 25x faster and 600x cheaper as an LLM judge — danshipper · 2026-09-16
- Robotics researcher pushes back on 'omni embodiment' hype: it's the hands, not the abstraction — chris_j_paxton · 2026-09-16
- Mathematician Tony Feng: current AI is not robustly superhuman yet — littmath · 2026-09-16
- Research team builds open science-task taxonomy from job postings — JMateosGarcia · 2026-09-16
- Anthropic researcher, aided by Claude, cracks Classic McEliece challenge instance — matthew_d_green · 2026-09-16
- StarVLA's VLAct trains VLAs on 16 GPUs by reshaping action representations, not data scaling — jiqizhixin · 2026-09-16