DeepSeekMath-V2 makes verification the product, scaling verifier compute ahead of the generator

le_james94 · x · 2026-09-16

Lecture 9 of a thread series argues DeepSeekMath-V2 treats verification as the product: train a verifier, then scale verification compute to stay ahead as the generator improves — though nobody knows if that holds outside math. Supporting evidence across the thread: METR puts o3's 50% time horizon at 110 minutes, doubling every 7 months; no system clears a 31% geometric mean on DeepScholar-Bench; AlphaCode 2 matches AlphaCode's 1M-sample performance with 100 samples (10,000x sample efficiency); DeepSeekMath's measurements show RLVR improves Maj@K but not Pass@K — more consistent, not fundamentally smarter; Math-Shepherd replaced 800K human labels with rollouts, lifting GSM8K from 77.9% to 84.1%.

Related event: Verifiers Take Center Stage as AI Benchmarks Split(2 posts)→

Original post →

More from Models

Models channel →