LastOPD: latent on-policy distillation collapses late, last-layer-only signal gains 5.55 on MATH-500

Jie Yang · hf · 2026-09-29

Distilling Qwen3-4B/8B into Qwen3-1.7B-Base, the authors find two counterintuitive failure modes in latent on-policy distillation (OPRD-style methods):

Analysis points to a mismatch: layers paired by depth play different roles in teacher and student, so continued alignment pulls the student toward teacher states it cannot understand.

The fix: LastOPD applies the latent signal only at the last-layer state — the common interface both LM heads read — and only during a 10-step crossfade into token-level OPD, preserving the useful signal before collapse sets in.

Results: +5.55 (4B teacher) and +4.02 (8B teacher) points on MATH-500 over token-only OPD, leads on most held-out datasets, and reaches token-only OPD's final score in about half the steps. Code is open-sourced.

Original post →

More from Models

Models channel →