Switch Distillation: Distill Only When the Teacher Is Confident to Keep Both Reasoning and Facts
StellaLisy · x · 2026-09-03
A new paper shows mid-training knowledge distillation boosts reasoning but hurts factual recall, since teachers are far less confident on knowledge-heavy tokens. Switch Distillation distills only at low teacher entropy, yielding up to 1.71x better reasoning and 1.19x better knowledge while preserving 97% factual recall.
Related event: Meta's Switch Distillation Fixes Mid-Training KD Trade-off(3 posts)→
More from Research
- WorldCrafter: a world model with persistent memory for consistent exploration — yshan2u · 2026-09-22
- Xiaomi MiMo-V2.6 details its largest RL scaling run: $2.6M, 1M-token contexts — KyeGomezB · 2026-09-22
- Microsoft's ShieldVLA Cuts VLA Safety Costs by 57% Using HJ Reachability — MicrosoftResearch · 2026-09-22
- UCSC's Prediction-Powered Smoothing Makes Disaggregated AI Evaluation More Precise — UCSantaCruz · 2026-09-22
- Rat-brain AI models land on AWS as biological computing goes mainstream — nordicinst · 2026-09-22
- OpenAI's unreleased model reportedly solved 100+ open math problems after 24 days of training — Confident_Salt_8108 · 2026-09-22