Deleting 90% of Weights: Song Han's Journey to Efficient AI & Quantization

JafarNajafov · x · 2026-08-05

A comprehensive thread details the groundbreaking contributions of MIT professor and NVIDIA research director Song Han in neural network compression, noting that almost every locally run quantized model today descends from his research.

The author concludes that the student who spent his PhD deleting weights now leads NVIDIA's Efficient AI team, forming the foundation for all local models running on consumer hardware.

Original post →

More from Infra

Infra channel →