DeepSeekMath data: RL on verifiable rewards improves Maj@K but not Pass@K
le_james94 · x · 2026-09-16
Pushing back on the claim that RL on verifiable rewards teaches models new capabilities, this post cites DeepSeekMath's own measurements: RL improves Maj@K but not Pass@K. The model became more consistent, not fundamentally smarter — part of a broader thread on verification-centric training like DeepSeekMath-V2.
Related event: Does RLVR Teach New Capabilities? Data Says Maybe Not(2 posts)→
More from Research
- Periodic Labs' Neon tops GPT-6 Astra on materials analysis using just 1,300 H200s — LiamFedus · 2026-09-16
- AI2's NGU sampling fixes RL for LLMs that only improves easy tasks — allenai · 2026-09-16
- CoLLAs 2026 orals: forgetting, sleep replay, and why LLMs can't play Hangman — apsarathchandar · 2026-09-16
- Jev Benchmark Launches: $42 per Billion Input Tokens, Output Free Forever — cephaloform · 2026-09-16
- Multi-agent RL post-training fights LLM mode collapse and boosts response diversity — natashajaques · 2026-09-16
- World Models Won't Get Us to AGI — Continual Learning Is the Missing Piece, and It's Hard — Intelligent-Cream-14 · 2026-09-16