4-bit KV Cache Tested: Perplexity Surges 43% in Long Contexts, q8_0 is the Sweet Spot

Dhan295 · reddit · 2026-08-02

A developer comprehensively tested KV cache quantization across different context lengths and found a counterintuitive phenomenon: 4-bit (q40) KV cache perplexity actually increases after 2K tokens, worsening by 43% at 8K.

Key Findings:

The author notes that testing was done on the smaller Qwen3-0.6B model, which is likely a worst-case scenario, but the conclusions sufficiently highlight the risks of blind KV quantization.

Original post →

More from Infra

Infra channel →