ETH Zurich: Repulsive Self-Distillation Destabilizes Training, Contrastive Distillation Wins
arkrause · x · 2026-10-02
Andreas Krause's group at ETH Zurich published a blogpost, On Repulsive and Attractive Teachers, dissecting how the direction of teacher signals shapes self-distillation.
- Following this year's popularity of on-policy distillation, some papers proposed flipping the self-distillation loss to move away from the teacher.
- Their findings: self-distillation with privileged context shifts student behavior, suppresses exploration, and can hurt reasoning; repulsive distillation reverses the shift but makes responses grow until hitting the cap, collapsing training.
- They propose contrastive self-distillation: pairing an attractive (correct-solution) teacher with a repulsive (incorrect-solution) teacher cancels behavioral shifts while preserving the correctness correction — better reasoning with stable training.
- The post includes clear token-level rescoring visualizations of all three schemes.
A high-quality methodological read for anyone following post-training and distillation.
More from Research
- JevBench adds multilingual queries to test Jev models across languages — airesearch12 · 2026-10-02
- Diffusion will be everywhere: why text diffusion models may replace autoregressive LLM inference — akbirthko · 2026-10-02
- BIABench: No AI agent scores above 0.19 on 3D bioimage analysis tasks — notredame · 2026-10-02
- BiasReducer from CMU edits only the reward head to adaptively cut length and confidence biases — CarnegieMellonU · 2026-10-02
- Anthropic's BootLoops: a toolkit for exact calculations in quantitative science — badumtsssst · 2026-10-02
- Jeff Clune keynote: open-ended and AI-generating algorithms will drive the AI science revolution — jeffclune · 2026-10-02