Two weeks of quant testing on a 5080: above Q4, most quants were statistically indistinguishable

KitchenAmoeba4438 · reddit · 2026-08-23

The author ran two straight weeks of quantization comparisons on an RTX 5080 (16GB) and concluded the opposite of community consensus: quantization isn't a continuous "quality dial" — most of the dial isn't connected.

Key findings:

Sessions were deliberately 2-4 messages, testing whether a quantized model is intact, not how it holds up under load. The next round targets DevOps, coding, and long sessions, where the author expects a VRAM/speed-vs-accuracy tradeoff rather than a simple ranking. Raw data, benchmark code, and datasets are open-sourced in the article's git repo.

Related event: Quantization below Q4 noticeably degrades local Agent models(3 posts)→

Original post →

More from Infra

Infra channel →