DeepSeekMath-V2 makes verification the product, scaling verifier compute ahead of the generator
le_james94 · x · 2026-09-16
Lecture 9 of a thread series argues DeepSeekMath-V2 treats verification as the product: train a verifier, then scale verification compute to stay ahead as the generator improves — though nobody knows if that holds outside math. Supporting evidence across the thread: METR puts o3's 50% time horizon at 110 minutes, doubling every 7 months; no system clears a 31% geometric mean on DeepScholar-Bench; AlphaCode 2 matches AlphaCode's 1M-sample performance with 100 samples (10,000x sample efficiency); DeepSeekMath's measurements show RLVR improves Maj@K but not Pass@K — more consistent, not fundamentally smarter; Math-Shepherd replaced 800K human labels with rollouts, lifting GSM8K from 77.9% to 84.1%.
Related event: Verifiers Take Center Stage as AI Benchmarks Split(2 posts)→
More from Models
- With retries and pooled selection, Qwen3.8 27B hits 92.04% on DeepSWE 1.1, ~18 pts above GPT-6 Astra — S_Conradi · 2026-09-16
- "Just output probability distributions, never hallucinate": AI safety claim gets mocked — inductionheads · 2026-09-16
- Chinese open models hit 53% of OpenRouter tokens, but closed models still dominate real adoption — ohlennart · 2026-09-16
- Niche AI Use Cases Keep Getting Absorbed Into General Models — samiramanabi · 2026-09-16
- Google reportedly building math-focused DeepThink variant, raw thoughts leak — PMinervini · 2026-09-16
- Player claims to find OpenAI GPT-6 "Astra" easter egg in Fallout 3 — imjustnewatai · 2026-09-16