Impact of KV Quantization on Qwen3.6-27B Tested
BitGreen1270 · reddit · 2026-07-08
The author systematically tested the impact of different KV quantization combinations on KL Divergence (KLD) for Qwen3.6-27B (bartowski quantization) using an RTX 5090. Key findings: Q8 generally outperforms Q6 and Q5; performance for Q8 and Q6 degrades sharply if the value uses q40; and Q5 is more tolerant of value quantization. The final recommendation is to use the highest quantization your VRAM allows and stick to the (q80, q80) KV quantization configuration, which results in almost no performance loss.
More from Infra
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Vera CPU is speeding up next-gen CPU and GPU design cycles — nordicinst · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27