Ant Group's PMOPD tackles capability seesaw in multi-teacher distillation
antgroup · hf · 2026-09-30
Ant Group's paper PMOPD addresses the "capability seesaw" in multi-teacher on-policy distillation (MOPD), where improving one domain suppresses others. The authors observe that parameter updates from different tasks rapidly concentrate in their own low-dimensional subspaces, providing a geometric basis for controlling cross-task interference.
PMOPD builds subspace memories from cumulative parameter displacements and projects gradients and optimizer updates to remove components interfering with protected task directions. A lightweight conflict probe characterizes task interactions and guides task ordering, plus a cycling strategy balancing subspace estimation with task revisitation. Across Code, Reasoning and Math tasks, PMOPD beats MOPD on every capability, lifting average scores by 2.54 points on Qwen2.5-7B and 2.09 on Llama-3.1-8B.
More from Research
- OpenAI's Navier-Stokes proof compiles clean, but the fluid vaporizes at 0.7 nm — OrganizationTop9026 · 2026-10-01
- Lancet Digital Health proposes clinician–AI scientist career pathway in medicine — pearsekeane · 2026-10-01
- Intent Lab's review agent catches 93% of 284 real bugs at 63% lower cost — jiayq · 2026-10-01
- RICE turns any LLM into a dense retriever with zero training — CShorten30 · 2026-10-01
- Functional ultrasound imaging: a computational deep dive into the new imaging modality — patrickmineault · 2026-10-01
- Three frontier models graded 50 tech leaders' public speech — LeCun ranks #1 — tonygwu · 2026-10-01