New Paper: Adding RL After OPD Consistently Beats Pure OPD, Pure RLVR, and Joint Methods
gregd_nlp · x · 2026-09-16
A new paper examines the interplay between OPD and RLVR: don't skip RL after OPD. Adding an RL phase after OPD-based reasoning training consistently improves performance, and the OPD→RL pipeline outperforms pure OPD, pure RLVR, and many joint OPD+RL methods across reasoning tasks—offering empirical guidance on training order.
Related event: Paper Finds OPD Followed by RL Beats Pure OPD or RLVR(3 posts)→
More from Research
- Single-axis random features beat full weight sets; a Markov-chain POS tagger trains in an hour — cephaloform · 2026-09-16
- Odyssey unveils Odyssey-3 foundation world model for robots, humanoids and driving — chris_j_paxton · 2026-09-16
- Jason Wei: Wet-Lab Data Lets Specialized Models Beat GPT-6 Astra at Frontier Science — vwxyzjn · 2026-09-16
- Hand-computable AI? Markov chains plus polynomial feature fits for POS tagging — cephaloform · 2026-09-16
- Pichai maps 9B genetic variants, launches WeatherNext 3 in Google's AI-for-science push — sundarpichai · 2026-09-16
- Most benchmark 'model errors' are actually benchmark or grading errors, analysis finds — zainhas · 2026-09-16