Impact of KV Quantization on Qwen3.6-27B Tested
BitGreen1270 · reddit · 2026-07-08
The author systematically tested the impact of different KV quantization combinations on KL Divergence (KLD) for Qwen3.6-27B (bartowski quantization) using an RTX 5090. Key findings: Q8 generally outperforms Q6 and Q5; performance for Q8 and Q6 degrades sharply if the value uses q40; and Q5 is more tolerant of value quantization. The final recommendation is to use the highest quantization your VRAM allows and stick to the (q80, q80) KV quantization configuration, which results in almost no performance loss.
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11