Same Q4KM label, sizes vary 150%: llama.cpp quant naming chaos called out

ParaboloidalCrest · reddit · 2026-09-03

A local inference user calls out the llama.cpp ecosystem over meaningless quantization labels: for one model, 'Q4KM' quants from different quantizers range from 94.5GB (AtomicChat) to 135GB (AesSedai), with lmstudio at 119GB and bartowski at 120GB — about 150% variance under the same name.

He argues the labels have become apples-to-oranges comparisons and is considering going back to making his own simple, honest Q40 quants. The post highlights a longstanding pain point: no shared standard for quant naming or quality benchmarking in the local model community.

Original post →

More from Infra

Infra channel →