DeepSeekMath data: RL on verifiable rewards improves Maj@K but not Pass@K

le_james94 · x · 2026-09-16

Pushing back on the claim that RL on verifiable rewards teaches models new capabilities, this post cites DeepSeekMath's own measurements: RL improves Maj@K but not Pass@K. The model became more consistent, not fundamentally smarter — part of a broader thread on verification-centric training like DeepSeekMath-V2.

Related event: Does RLVR Teach New Capabilities? Data Says Maybe Not(2 posts)→

Original post →

More from Research

Research channel →