Direct-OPD Transfers Weak Model RL Gains to Strong Models
tokenbender · x · 2026-08-21
The paper proposes Direct On-Policy Distillation (Direct-OPD) to avoid expensive RL runs on large models. It runs RL on a weaker model and transfers the policy shift as an implicit reward to the stronger student. Empirically, Direct-OPD consistently improves stronger target models, boosting Qwen3-1.7B accuracy from 48.3% to 58.3%.
More from Models
- Mystery Model Ox Alpha Review: Near-Fable Performance, Unique Vision/Logic Profile — Afinetheorem · 2026-08-22
- Funny Fail: Asking OpenAI for a Hotel Suggests Founding Expedia — dbasch · 2026-08-22
- New NInfer and oQ4e Quants for Ornith 1.5 and Qwen 3.8 — Pyros-SD-Models · 2026-08-22
- Rumor: Zhipu's GLM 5.3 Flash and Kimi k3.1 incoming on OpenRouter — calabi_and_yau · 2026-08-21
- Users report ChatGPT's memory has noticeably degraded over the past month — Middle-Wealth-6755 · 2026-08-21
- Chinese stealth model Ox Alpha reviewed: 1M context but prone to spinning — bindureddy · 2026-08-21