Deleting 90% of Weights: Song Han's Journey to Efficient AI & Quantization
JafarNajafov · x · 2026-08-05
A comprehensive thread details the groundbreaking contributions of MIT professor and NVIDIA research director Song Han in neural network compression, noting that almost every locally run quantized model today descends from his research.
- Early Breakthroughs: During his PhD at Stanford in 2015, he developed 'Deep Compression'—combining pruning, quantization, and Huffman coding to shrink ResNet-50 from 100MB to 6MB with no accuracy loss, winning ICLR 2016 Best Paper.
- Hardware & Ecosystem: His sparse architecture EIE directly inspired NVIDIA's Sparse Tensor Cores in the 2020 Ampere architecture. He co-founded DeePhi Tech (acquired by Xilinx/AMD) and OmniML (acquired by NVIDIA).
- The LLM Era: As LLMs hit compute walls, his lab shipped SmoothQuant (fixing 8-bit PTQ breakdowns) and AWQ. AWQ protects a tiny fraction of salient weights to eliminate most quantization error without backprop, won MLSys 2024 Best Paper, and was adopted natively by NVIDIA, Google, and HuggingFace to put 70B models on mobile GPUs.
The author concludes that the student who spent his PhD deleting weights now leads NVIDIA's Efficient AI team, forming the foundation for all local models running on consumer hardware.
More from Infra
- Ollama Auto-Enables Qwen3.5 MTP on Macs, MLX Backend Shows Major Speedup — BTA_Labs · 2026-08-05
- Performance Deep-Dive: Numpy and CPython in the Free-Threaded Build — abhi9u · 2026-08-05
- Discussing Groq's Ultra-Fast Inference: Does Real-World Speed Compromise Quality? — Scared-Tip7914 · 2026-08-05
- xAI Supercomputer in Memphis Coincides with 65% Surge in Housing Inventory — brianrkelly · 2026-08-05
- Dell's Son Raises $1B for BasePower, Valuing Home Battery Startup at $13B — 创业邦 · 2026-08-05
- Qwen3.6 35B NVFP4 Hits 3,000 Tokens/s on a Single RTX PRO 6000 — max_paperclips · 2026-08-05