Paper Finds OPD Followed by RL Beats Pure OPD or RLVR
An EMNLP 2026 Findings paper shows that on-policy distillation followed by an RL stage consistently outperforms pure OPD, pure RLVR, and joint optimization across reasoning tasks.
2026-09-16 ~ 2026-09-16 · 3 related posts
- EMNLP Paper: Simple OPD-then-RLVR Two-Stage Training Beats All Joint Distillation+RL Baselines — xiye_nlp · 2026-09-16
- New Paper: Adding RL After OPD Consistently Beats Pure OPD, Pure RLVR, and Joint Methods — gregd_nlp · 2026-09-16
- Sequential Beats Joint: OPD-then-RLVR two-stage training consistently outperforms joint optimization — xiye_nlp · 2026-09-16