LastOPD: latent on-policy distillation collapses late, last-layer-only signal gains 5.55 on MATH-500
Jie Yang · hf · 2026-09-29
Distilling Qwen3-4B/8B into Qwen3-1.7B-Base, the authors find two counterintuitive failure modes in latent on-policy distillation (OPRD-style methods):
- Early gain, late collapse: latent supervision alone lifts MATH-500 from 25 to 46, but continued training degrades it to 11 with no recovery;
- Better alignment, worse behavior: the alignment metric keeps improving through the collapse, and the most aligned model performs worst.
Analysis points to a mismatch: layers paired by depth play different roles in teacher and student, so continued alignment pulls the student toward teacher states it cannot understand.
The fix: LastOPD applies the latent signal only at the last-layer state — the common interface both LM heads read — and only during a 10-step crossfade into token-level OPD, preserving the useful signal before collapse sets in.
Results: +5.55 (4B teacher) and +4.02 (8B teacher) points on MATH-500 over token-only OPD, leads on most held-out datasets, and reaches token-only OPD's final score in about half the steps. Code is open-sourced.
More from Models
- Small Korean Team Nemotron Labs Hits Top 5 in Open Model Rankings — NVIDIAAI · 2026-09-29
- OpenAI reportedly delayed GPT OSS over Kimi K2 competition, not safety — johnohallman · 2026-09-29
- LLM shows surprisingly usable calibration classifying abstracts on human subjects — RexDouglass · 2026-09-29
- Dev flags suspected Opus 5.5 hardcoding of 'humans are right, AIs are wrong' bias — repligate · 2026-09-29
- OpenAI Says It Will Not Release Newest AI Model Over Safety Concerns — jbegley · 2026-09-29
- Speculation: Meta paid full API prices for Fable traces to distill, and outputs taste like Claude — andersonbcdefg · 2026-09-29