A Quantization Handbook: From Affine Quantization to GPTQ, AWQ, QLoRA and FP8
techNmak · x · 2026-09-28
techNmak published a systematic handbook on how quantization actually works, aiming to explain what "4-bit" really means and why fewer bits save memory without automatically speeding up inference.
Coverage includes:
- Affine integer quantization basics: scales, zero-points, rounding, clipping, and granularities
- Symmetric/asymmetric schemes, per-tensor/per-channel/per-group scaling, calibration, PTQ vs QAT
- Weight-only and activation quantization, outlier handling
- GPTQ, AWQ, SmoothQuant, LLM.int8(), NF4, QLoRA
- FP8 and KV-cache quantization (KIVI, KVQuant)
- Kernel/hardware considerations that make low precision actually useful
Related event: Comprehensive Handbook on LLM Quantization Released(2 posts)→
More from Infra
- Running 8 watercooled GPUs for local AI: one user's case for watercooling over air cooling — HanchungLee · 2026-09-28
- HF: transformers backend now matches native vLLM speed, no porting needed — ariG23498 · 2026-09-28
- Renting their cluster's compute would have cost over $1 billion on a 5-year deal — ericzelikman · 2026-09-28
- Self-Hosted K3 Cluster Lets Autonomous Infra Research Burn Tokens Without Worry — Xianbao_QIAN · 2026-09-28
- ByteDance Seed Proposes PISA: Block-Sparse Attention With O(Nlog N) Complexity — ByteDance-Seed · 2026-09-28
- Cranking llama.cpp for one model: Reddit proposal claims 2x+ inference gains — segmond · 2026-09-28