Sequential Beats Joint: OPD-then-RLVR two-stage training consistently outperforms joint optimization

xiye_nlp · x · 2026-09-16

An EMNLP 2026 Findings paper shows a simple two-stage scheme — on-policy distillation (OPD) followed by RLVR — consistently beats pure OPD, pure RLVR, and all joint baselines (weighted-additive or teacher-modulated) on logic and math benchmarks. The mechanism: OPD expands the student's coverage of teacher-supported solutions, RL sharpens within that support, and jointly optimizing the two signals causes interference. Practical recipe: switch when the OPD validation score plateaus, and OPD is a better RL cold start than SFT. After switching, top-K overlap drops while probability mass on shared tokens rises, showing RL sharpens within the teacher's support.

Related event: Paper Finds OPD Followed by RL Beats Pure OPD or RLVR(3 posts)→

Original post →

More from Research

Research channel →