Why distillation beats vanilla supervised learning: it passes the full probability vector
khademinori · x · 2026-09-29
A concise ML insight on knowledge distillation: unlike vanilla supervised learning, which communicates a single one-hot label, distillation transfers the entire probability distribution vector p() from the teacher to the student. This "dark knowledge" encodes inter-class similarity, making the training signal much richer and explaining why distillation is unusually effective.
More from Research
- HalluWorld Benchmark, Accepted at NeurIPS, Shows Models Nail Perception but Fail Simulation — xennygrimmato_ · 2026-09-29
- Quintin Pope: backdoor-style training setups are a poor stand-in for hypothesized inner optimizers — QuintinPope5 · 2026-09-29
- Ocular microtremors at ~100Hz may give the visual cortex apparent super-resolution — docmilanfar · 2026-09-29
- AI-designed viruses from scratch: 302 phage genomes synthesized, 16 worked — CallRevolutionary894 · 2026-09-29
- Reddit thread: softmax has only N-1 degrees of freedom — drop one input? — Kinexity · 2026-09-29
- EAPO: entropy-guided credit assignment for RLVR improves exploration in LLM reasoning — coallaoh · 2026-09-29