Same Q4KM label, sizes vary 150%: llama.cpp quant naming chaos called out
ParaboloidalCrest · reddit · 2026-09-03
A local inference user calls out the llama.cpp ecosystem over meaningless quantization labels: for one model, 'Q4KM' quants from different quantizers range from 94.5GB (AtomicChat) to 135GB (AesSedai), with lmstudio at 119GB and bartowski at 120GB — about 150% variance under the same name.
He argues the labels have become apples-to-oranges comparisons and is considering going back to making his own simple, honest Q40 quants. The post highlights a longstanding pain point: no shared standard for quant naming or quality benchmarking in the local model community.
More from Infra
- Blogger argues Zeiss's mirror-coating know-how is only ~30 years deep, not magic — teortaxesTex · 2026-09-03
- Tenstorrent and AI & Inc launch JapanFold: free inference for open-source drug discovery models — DavidBennett__ · 2026-09-03
- MediaTek may take Google's 1M training TPU business as Broadcom keeps inference — firstadopter · 2026-09-03
- Voice as default input: local STT with Mac Parakeet, Whisper vs Parakeet tradeoffs — LeatherRub7248 · 2026-09-03
- How much VRAM/RAM is a meaningful upgrade for local LLM inference? Reddit weighs in — MiceLiceandVice · 2026-09-03
- Zeiss says China is still ~15 years from EUV lithography; skeptics doubt supplier's word — basedjensen · 2026-09-03