Latent-MOPD: Representation-Level Multi-Teacher Distillation Lets One Student Beat Every Teacher
zillow · hf · 2026-10-06
Latent-MOPD is claimed to be the first representation-level multi-teacher on-policy distillation method for LLMs, transferring knowledge through both output distributions and hidden states without extra teacher training.
- Late-layer targets, a shared projection for unequal widths, and domain-grouped updates coordinate supervision; each teacher's signal shifts from hidden states to token predictions over training
- In the same-family setting it beats token-only, representation-only, and uniform-averaging baselines on all nine math/code/logic benchmarks; a student matched to each teacher's size surpasses the per-benchmark best teacher on most benchmarks
- With larger cross-family teachers it outperforms both single-channel baselines everywhere
- Ablations show all-layer representation-only training stays stable with domain-pure updates but collapses when teacher domains are interleaved
More from Research
- 1-in-4 to 1-in-3 dissertations now written with AI, replicated study finds — skdh · 2026-10-06
- Researcher maps AI Village data to Ising and subcritical Hawkes systems — ctjlewis · 2026-10-06
- Poisoned Conversation: Privacy-Leaking Watermarks hit 100% TPR in unified multimodal models — chaumian · 2026-10-06
- Truly subcubic APSP and truly subquadratic 3SUM breakthrough reshapes fine-grained complexity — thegautamkamath · 2026-10-06
- Amazon's ALoDLM: Token-Adaptive Looped Diffusion LMs Beat AR Baselines — amazon · 2026-10-06
- CANOPY: Adaptive Evidence Compression Cuts Multimodal RAG Tokens Up to 28% — POSTECH · 2026-10-06