Why the Same Llama 3.2 1B Comes in Different File Sizes: Quantization Explained
night-alien · reddit · 2026-09-29
A beginner-friendly explainer on why the same Llama 3.2 1B model ships in different file sizes. The key idea is quantization: compressing weights from high-precision floats to lower-bit integers to shrink files and speed up inference.
The article covers what Q2, Q4, Q8, and F16 mean — the number of bits per weight — and the trade-off between smaller files and precision loss. The author invites feedback and corrections.
More from Research
- Amazon's GEB grounds entity biographies for long-video memory, hits 72% on EgoLifeQA — amazon · 2026-09-30
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- CrossBFM Distills a Shared Behavior Space Across Humanoid Robots in Under One GPU-Hour — Tan-Dzung Do · 2026-09-30
- AnisoWM: anisotropic representations improve planning in JEPA world models — SeoulNatlUniv · 2026-09-30
- NVIDIA's HDL Cuts RLVR Token Cost 2.5x by Localizing Where Models Change Their Mind — nvidia · 2026-09-30
- Meituan's SAKI: Coupling-Routed Teacher Supervision Speeds Up On-Policy Distillation 4.22x — meituan · 2026-09-30