Δ-MOPD transfers teacher shifts instead of endpoints, +4.11 Math in multi-teacher distillation

LinkedIn · hf · 2026-10-09

LinkedIn releases Δ-MOPD, improving multi-teacher on-policy distillation (MOPD). Endpoint supervision transfers each teacher's final policy but mixes in preferences inherited from the teacher's base; the paper shows inherited base pull can exceed the post-training shift, impeding transfer.

Δ-MOPD instead transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization:

Target construction is an independent design axis in MOPD, complementary to teacher selection.

Original post →

More from Research

Research channel →