New Paper: Adding RL After OPD Consistently Beats Pure OPD, Pure RLVR, and Joint Methods

gregd_nlp · x · 2026-09-16

A new paper examines the interplay between OPD and RLVR: don't skip RL after OPD. Adding an RL phase after OPD-based reasoning training consistently improves performance, and the OPD→RL pipeline outperforms pure OPD, pure RLVR, and many joint OPD+RL methods across reasoning tasks—offering empirical guidance on training order.

Related event: Paper Finds OPD Followed by RL Beats Pure OPD or RLVR(3 posts)→

Original post →

More from Research

Research channel →