Does RLVR Teach New Capabilities? Data Says Maybe Not
DeepSeekMath's own data shows RL improves Maj@K but not Pass@K, challenging claims that RLVR teaches new capabilities; related discussion frames feedback signals as a ladder from environment answers to constitutional AI.
2026-09-16 ~ 2026-09-16 · 2 related posts
- A ladder of weaker feedback signals: ReAct, RLEF, Constitutional AI — and Math-Shepherd's rollouts — le_james94 · 2026-09-16
- DeepSeekMath data: RL on verifiable rewards improves Maj@K but not Pass@K — le_james94 · 2026-09-16