Paper Finds OPD Followed by RL Beats Pure OPD or RLVR

An EMNLP 2026 Findings paper shows that on-policy distillation followed by an RL stage consistently outperforms pure OPD, pure RLVR, and joint optimization across reasoning tasks.

2026-09-16 ~ 2026-09-16 · 3 related posts