EMNLP Paper: Simple OPD-then-RLVR Two-Stage Training Beats All Joint Distillation+RL Baselines
xiye_nlp · x · 2026-09-16
An EMNLP 2026 Findings paper (arXiv:2609.04108) by Boyan Li, Bingsen Chen et al. shows that a simple two-stage scheme—on-policy distillation (OPD) followed by RLVR—consistently outperforms pure OPD, pure RLVR, and all joint baselines (weighted-additive, teacher-modulated rescaling) on logic and math reasoning benchmarks.
Key findings:
- Mechanism: OPD expands the student's coverage of teacher-supported solutions; RL sharpens within that support. Optimizing both simultaneously causes interference. After switching, top-K overlap drops while probability mass on shared tokens rises.
- Practical recipe: the OPD validation score is the key signal for when to switch—run OPD until validation plateaus, then switch to RLVR. No new objective needed.
- OPD is a better cold start for RL than SFT.
Paper and code are open-sourced.
Related event: Paper Finds OPD Followed by RL Beats Pure OPD or RLVR(3 posts)→
More from Research
- Therna Bio opens Chronos platform, releases gene expression datasets with millions of measurements — AllThingsApx · 2026-09-16
- GlossoGen: open-source platform shows AI agents spontaneously inventing new languages under pressure — EliasEskin · 2026-09-16
- Paper: upsampling alignment discourse in pretraining cuts misalignment from 45% to 9% — TuhinChakr · 2026-09-16
- One researcher taught world model Odyssey-3 to drive in India with just 20 hours of data — soleio · 2026-09-16
- Is there a real hardware lottery in AI? Transformers may have won by fitting GPUs — prateekj · 2026-09-16
- Anthropic: automated Claude alignment researchers mitigate 10 failure classes, generalize to 4.7x larger models — echen · 2026-09-16