Multi-domain on-policy distillation lets one model learn many skills at once
helloiamleonie · x · 2026-09-25
Leonie explains MOPD (multi-domain on-policy distillation) used in LFM2.5 training: extending on-policy distillation to multiple teachers so the student model learns different skills simultaneously, rather than sequentially. It's the key step that merges specialized teachers' abilities back into a single checkpoint.
Related event: Liquid AI Releases LFM2.5-2.6B and Details Its On-Device Training Recipe(10 posts)→
More from Research
- Nature paper: psilocybin reshapes latent temporal structure in brain activity — adeelrazi · 2026-09-25
- Francis Bach Blog: Taming Exploding Variance of Exponential Means With Least Squares — BachFrancis · 2026-09-25
- William & Mary AI Frontier Lab lands 5 NeurIPS acceptances — jindong_wang92 · 2026-09-25
- RUC open-sources EvoOntology, a self-evolving ontology layer for data agents via MCP — JeremyCMorgan · 2026-09-25
- Inside LFM2.5-2.6B's post-training recipe: SFT, RL, multi-domain distillation — helloiamleonie · 2026-09-25
- Transferring Qwen3.8's n-gram memory into a 0.8B model cuts perplexity 5.05% — Nicolodeva · 2026-09-25