Cal-OPD: Calibrated on-policy distillation beats standard OPD with half the signal

nanjinguniv · hf · 2026-09-21

This paper introduces Calibrated On-Policy Distillation (Cal-OPD). Standard on-policy distillation learns the token-level teacher–student discrepancy, but this signal mixes in the teacher's own deviations, exacerbated by privileged OPD's larger teacher-side likelihood shifts. Cal-OPD estimates the teacher's self-deviation region via positive and negative privileged interventions and keeps only the discrepancy beyond it. On math reasoning benchmarks, using only 52–65% of the original discrepancy as training signal, Cal-OPD consistently outperforms standard OPD and variants across model scales.

Original post →

More from Research

Research channel →