Two weeks of quant testing on a 5080: above Q4, most quants were statistically indistinguishable
KitchenAmoeba4438 · reddit · 2026-08-23
The author ran two straight weeks of quantization comparisons on an RTX 5080 (16GB) and concluded the opposite of community consensus: quantization isn't a continuous "quality dial" — most of the dial isn't connected.
Key findings:
- Damage starts below Q4; most quants above it were statistically indistinguishable from each other under these tests
- QAT breaks on mismatch: a model QAT-trained at one quant run at another performs badly — match the trained quant or pick a different model
- MoEs are less affected than dense models
- Gaps between common quants were smaller than consensus predicts; the author went in expecting to confirm consensus and the data didn't support it
Sessions were deliberately 2-4 messages, testing whether a quantized model is intact, not how it holds up under load. The next round targets DevOps, coding, and long sessions, where the author expects a VRAM/speed-vs-accuracy tradeoff rather than a simple ranking. Raw data, benchmark code, and datasets are open-sourced in the article's git repo.
Related event: Quantization below Q4 noticeably degrades local Agent models(3 posts)→
More from Infra
- SOTA compute rack's idle power matches a blue whale at 1/100th the mass — jwt0625 · 2026-08-24
- Samsung Details Custom HBM Base Die Advantages: 7 Key Benefits of Leading Edge Nodes — BenBajarin · 2026-08-24
- Custom HBM Base Die Becoming Competitive Vector for Memory Giants — BenBajarin · 2026-08-24
- Samsung presenter criticizes Micron's HBM4 base die design — zephyr_z9 · 2026-08-24
- Can a vehicle wrap hide your car from Flock cameras? — HankYeomans · 2026-08-24
- Feynman Ultra may hit 100 TB/s memory bandwidth via HBM5 — zephyr_z9 · 2026-08-24