Sequential Beats Joint: OPD-then-RLVR two-stage training consistently outperforms joint optimization
xiye_nlp · x · 2026-09-16
An EMNLP 2026 Findings paper shows a simple two-stage scheme — on-policy distillation (OPD) followed by RLVR — consistently beats pure OPD, pure RLVR, and all joint baselines (weighted-additive or teacher-modulated) on logic and math benchmarks. The mechanism: OPD expands the student's coverage of teacher-supported solutions, RL sharpens within that support, and jointly optimizing the two signals causes interference. Practical recipe: switch when the OPD validation score plateaus, and OPD is a better RL cold start than SFT. After switching, top-K overlap drops while probability mass on shared tokens rises, showing RL sharpens within the teacher's support.
Related event: Paper Finds OPD Followed by RL Beats Pure OPD or RLVR(3 posts)→
More from Research
- AgentIR embeds agent reasoning traces into retrieval, plus a BrowseComp-Plus benchmark — CShorten30 · 2026-09-16
- AI Evals FAQ grows to 48 Q&As: sensitive data, huge traces, and stale gold datasets — HamelHusain · 2026-09-16
- TurnTrout Offers Shard Theory Explanation for Assistant Behaviors Shaped by AI-Free Documents — dhadfieldmenell · 2026-09-16
- NVIDIA releases FoundationPose on Hugging Face: unified 6-DoF pose model, no fine-tuning needed — _akhaliq · 2026-09-16
- 3 weeks through Stanford CS329A: the generator has outrun the verifier — le_james94 · 2026-09-16
- DeepSeekMath-V2 makes verification the product, scaling verifier compute ahead of the generator — le_james94 · 2026-09-16