Audit reveals 64 GGUF quants mislabeled across 25 repos
Daxfortuna · reddit · 2026-08-29
An audit of 443 GGUF quantized models across 25 repos found that 64 have actual quantization types that differ from their filenames.
The Issue:
K-quants and i-quants require the first tensor dimension to be divisible by 256. When this condition isn't met, llama-quantize silently falls back to a compatible type (e.g., IQ4NL or Q40). This means models labeled as extremely low-bit (e.g., IQ2XXS) may actually have a bit-per-weight (bpw) closer to 4.5, offering far less compression than advertised.
Affected Models:
- Nemotron-3.5-Lightning: 99% of parameters forced into fallback. Different labeled rungs (2.06-2.56 bpw) all measured at 4.58 bpw.
- Qwen3.8-Flash-Next: 51.9% of parameters forced into fallback.
- Nemotron-3-Super-120B: 18 of 23 quant rungs contain fallbacks.
Clean Results:
MiniMax-M2.1, byteshape's Qwen3.6 quants, and bartowski's Ornith-1.5 showed zero forced tensors with accurate labeling.
Recommendation: When downloading low-bit quants where filenames cannot guarantee accuracy, it may be better to choose honestly labeled Q40 or IQ4NL rather than chasing advertised compression ratios.
More from Infra
- Conifer SDK Open-Sourced: Unified Gateway with Exact Cost Tracking — ycombinator · 2026-08-29
- MiniMax H3 Max sets new video generation speed frontier, 24x faster — JenniferHli · 2026-08-29
- Ruff adopts PGO, boosting performance by 8-15% — charliermarsh · 2026-08-29
- MiniMax Releases Open Fast H3 v1: 14x Video Gen Speedup — ArashVahdat · 2026-08-29
- Prediction market: 76% chance Nvidia remains largest company by year-end — Polymarket · 2026-08-29
- FreeToken Engine: Run 35B Models at 39.3 t/s on 8GB GPUs — rohanpaul_ai · 2026-08-29