8-bit KV Cache Severely Drags Down Local Inference Performance

Sentdex · x · 2026-07-03

Sentdex discovered a counterintuitive phenomenon: many assume an 8-bit KV cache is "free" for local context, but testing GLM 5.2 (4-bit weights with an 8-bit KV) caused a massive performance hit—Terminal Bench scores plummeted from 46/89 to 21/89. The conclusion is that using a 2-bit version of GLM 5.2 is better than resorting to an 8-bit KV cache.

Original post →

More from Infra

Infra channel →