RetireOPD: self-retiring on-policy distillation lifts agent RL success rates by up to 18.8%
Yan Yu · hf · 2026-09-18
RetireOPD addresses sparse rewards in multi-turn agent RL by first training a decoupled, skill-conditioned teacher with environment rewards, then training a skill-free student with joint RL and on-policy distillation. Instead of a fixed schedule, Adaptive Retirement lets the student drop the teacher once their discrepancy plateaus and it reaches a target fraction of the teacher's success rate. On Qwen2.5 models (1.5B–7B), it improves ALFWorld success by 14.1–18.8 points and WebShop accuracy by 11.8–19.0 points over RL baselines, surpassing its own teacher in every setting.
More from Research
- ML is leaving its alchemy era: assumed truths can now be instrumented and falsified — mike64_t · 2026-09-18
- Together AI paper: score centering cancels drift in off-policy RL under train-inference mismatch — PandaAshwinee · 2026-09-18
- Cell's AI-in-biology special issue debuts LongevityBench, a 17-task aging benchmark — JosephJacks_ · 2026-09-18
- After OpenAI cracks Navier-Stokes, the era of the human mathematician is ending — yeastsplainer · 2026-09-18
- Custom BB-SLAM hits 0.02m accuracy with pure visual odometry, no LiDAR or IMU — broodsugar · 2026-09-18
- AI4Science Breakthrough Said to Exceed Expectations, Could Transform Computational Chemistry — BenBlaiszik · 2026-09-18