Distillation criticized: gradient descent can't even copy the teacher model
deliprao · x · 2026-09-26
- Delip Rao argues that building a Jev-like model via distillation hurts you: citing Andrew Wilson's team's NeurIPS 2021 paper, gradient descent often fails to find parameters that copy the teacher even when they exist and you start nearby — and students systematically exaggerate the teacher's overconfidence.
- The paper shows teacher-student predictive distributions often remain far apart even with sufficient student capacity, and matching the teacher more closely doesn't always improve generalization.
- Takeaway: users can sniff out benchmaxxed models; doing the right thing pays off.
More from Research
- Anthropic: Claude computes nine-loop scattering amplitudes, breaking the eight-loop record — LucaAmb · 2026-09-26
- Microsoft open-sources Fabric-RLM: LLMs write code to recursively chew through big data — adnan_hashmi · 2026-09-26
- New paper asks where to draw the line on mental privacy as BCI decoding improves — melnykowycz · 2026-09-26
- Where to get the 277-page Foundations of LLMs PDF and what chapter 5 covers — mdancho84 · 2026-09-26
- LLM textbook thread: decoding algorithms, acceleration, and inference-time scaling in chapter 5 — mdancho84 · 2026-09-26
- Free 277-page LLM textbook Foundations of Large Language Models updated with a new chapter — mdancho84 · 2026-09-26