Relay-OPD improves on-policy distillation by handing failed prefixes back to the teacher
zju · hf · 2026-07-29
Relay-OPD addresses a prefix failure issue in on-policy distillation: when a student goes down the wrong reasoning path, later supervision becomes noisy and wastes compute.
- It detects failed-prefix handoff points with a label-free trigger based on teacher–student continuation asymmetry.
- During training, the teacher briefly takes over at those points to generate a relay trajectory, then the student resumes optimization on that trajectory.
- The method uses a limited relay budget to focus intervention on early critical positions while staying close to the student policy.
- With Qwen3-4B-Instruct-2507 as teacher and Qwen3-0.6B / 1.7B-Non-Thinking students, it achieves the best or second-best result on every one of 8 math reasoning benchmarks.
- Compared with standard OPD, it improves average performance by +5.73% for the 1.7B student and beats FastOPD by +1.49%; the training trajectory length is cut by over 50%.
More from Research
- 30 years on, deep learning still lacks a clear answer on flat minima — fleetwood___ · 2026-07-29
- Borrowing from Antiquity: New Framework Tackles 'Silent Failures' in Multi-Agent Chains — alizahidrajaa · 2026-07-29
- ISNAD brings claim-level provenance to multi-agent LLM chains — alizahidrajaa · 2026-07-29
- Goodfire says its block-sparse featurizers preserve the shape of model reasoning — bendee983 · 2026-07-29
- MODUS turns a decoder-only model into a single any-to-any multimodal system — EPFL-VILAB · 2026-07-29
- Manski says clinical statistics research has a systemic methodological dysfunction — RexDouglass · 2026-07-29