EMNLP Paper: Simple OPD-then-RLVR Two-Stage Training Beats All Joint Distillation+RL Baselines

xiye_nlp · x · 2026-09-16

An EMNLP 2026 Findings paper (arXiv:2609.04108) by Boyan Li, Bingsen Chen et al. shows that a simple two-stage scheme—on-policy distillation (OPD) followed by RLVR—consistently outperforms pure OPD, pure RLVR, and all joint baselines (weighted-additive, teacher-modulated rescaling) on logic and math reasoning benchmarks.

Key findings:

Paper and code are open-sourced.

Related event: Paper Finds OPD Followed by RL Beats Pure OPD or RLVR(3 posts)→

Original post →

More from Research

Research channel →