Cal-OPD: Calibrated on-policy distillation beats standard OPD with half the signal
nanjinguniv · hf · 2026-09-21
This paper introduces Calibrated On-Policy Distillation (Cal-OPD). Standard on-policy distillation learns the token-level teacher–student discrepancy, but this signal mixes in the teacher's own deviations, exacerbated by privileged OPD's larger teacher-side likelihood shifts. Cal-OPD estimates the teacher's self-deviation region via positive and negative privileged interventions and keeps only the discrepancy beyond it. On math reasoning benchmarks, using only 52–65% of the original discrepancy as training signal, Cal-OPD consistently outperforms standard OPD and variants across model scales.
More from Research
- AI writing detectors flagged her lab vision post; she asks what we should actually measure — furongh · 2026-09-21
- Human-AI Collaboration Settles Major Open Problem in Multi-Winner Voting Theory — xuanalogue · 2026-09-21
- Benchmarks show coding agents edit code they shouldn't in 35-65% of cases; prompt framing is the lever — RunAI_Coder · 2026-09-21
- Stanford and Arc Institute use AI to design 16 functional phages that kill resistant bacteria — emmanuelvivier · 2026-09-21
- EvalSeal v1.5.0: open-source reproducibility receipts for LLM evals — Fit_Fortune953 · 2026-09-21
- Program-as-Weights: 0.6B model matches Qwen3-32B prompting with 1/50th memory, runs locally — yuntiandeng · 2026-09-21