Δ-MOPD transfers teacher shifts instead of endpoints, +4.11 Math in multi-teacher distillation
LinkedIn · hf · 2026-10-09
LinkedIn releases Δ-MOPD, improving multi-teacher on-policy distillation (MOPD). Endpoint supervision transfers each teacher's final policy but mixes in preferences inherited from the teacher's base; the paper shows inherited base pull can exceed the post-training shift, impeding transfer.
Δ-MOPD instead transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization:
- With three composed teachers: +4.11 Math and +1.95 five-benchmark points over endpoint composition
- Under phased routing, the training-order gap drops from 10.50 to 6.42 points
- Under interleaved routing, both targets perform comparably
Target construction is an independent design axis in MOPD, complementary to teacher selection.
More from Research
- LLMs reportedly advance on 4 of 7 Millennium Prize Problems, 3hrs compute each — ycombinator · 2026-10-09
- New research: conflicting training values can make models' CoT contradict their answers — OwainEvans_UK · 2026-10-09
- Arena raises $200M Series B at $3.1B valuation, launches agent Alignment Index — a16z · 2026-10-09
- PAMI anchors object motion to body parts for text-to-HOI, +14.5% contact recall — Chuqiao Li · 2026-10-09
- Lineage-gated agent memory wipes 18.8-25.5% cross-department leakage at 13.8μs overhead — Venkata M Sangaraju · 2026-10-09
- When should a fast agent defer? Deferring least-confident 30% to a reasoner gains up to +0.13 — Gian Luca Bailo · 2026-10-09