Solving Privilege Illusion in Policy Distillation: DAPD Boosts Model Performance

Shanghai-AI-Laboratory · hf · 2026-08-04

On-policy self-distillation (OPSD) is widely used for language-model post-training. However, it can induce a 'privilege illusion': the student learns privilege-dependent behavior from the teacher during training, but behaves as if the privileged information is still available during inference, ultimately degrading performance.

Researchers from Shanghai AI Lab identify information asymmetry between the privileged teacher and the student at inference as the root cause. They propose DAPD (Dual-Anchored Policy Distillation):

Experiments show DAPD outperforms OPSD by +2.00 points on average across tasks, with gains persisting across scales from 4B to 32B models.

Original post →

More from Research

Research channel →