Why distillation beats vanilla supervised learning: it passes the full probability vector

khademinori · x · 2026-09-29

A concise ML insight on knowledge distillation: unlike vanilla supervised learning, which communicates a single one-hot label, distillation transfers the entire probability distribution vector p() from the teacher to the student. This "dark knowledge" encodes inter-class similarity, making the training signal much richer and explaining why distillation is unusually effective.

Original post →

More from Research

Research channel →