Ant Group's PMOPD tackles capability seesaw in multi-teacher distillation

antgroup · hf · 2026-09-30

Ant Group's paper PMOPD addresses the "capability seesaw" in multi-teacher on-policy distillation (MOPD), where improving one domain suppresses others. The authors observe that parameter updates from different tasks rapidly concentrate in their own low-dimensional subspaces, providing a geometric basis for controlling cross-task interference.

PMOPD builds subspace memories from cumulative parameter displacements and projects gradients and optimizer updates to remove components interfering with protected task directions. A lightweight conflict probe characterizes task interactions and guides task ordering, plus a cycling strategy balancing subspace estimation with task revisitation. Across Code, Reasoning and Math tasks, PMOPD beats MOPD on every capability, lifting average scores by 2.54 points on Qwen2.5-7B and 2.09 on Llama-3.1-8B.

Original post →

More from Research

Research channel →