Meta research: Knowledge distillation cuts training data memorization by over 50% vs fine-tuning

lasha_nlp · x · 2026-10-02

Jaydeep Borkar will present this Meta memorization study at COLM (Oct 6, Grand Ballroom, poster #108) in San Francisco. Key finding: knowledge distillation doesn't just improve LLM performance—it also makes models memorize substantially less training data than standard fine-tuning, reducing memorization by more than 50%. That makes distillation a practical lever for lowering privacy and copyright leakage risks in training data. The author is also open to discussing memorization, safety, synthetic data pretraining, and is entering the industry job market this fall.

Original post →

More from Research

Research channel →