UCLA paper gives KV cache eviction a mathematical foundation via importance sampling

burkov · x · 2026-09-03

A UCLA paper puts theory behind the common practice of pruning LLM KV caches. It shows picking optimal entries is computationally hard, then rewrites attention as an expected value and treats eviction as estimating it: entries are sampled by estimated importance, and importance sampling corrects the attention computation for what was removed, giving an explicit error account for cache eviction.

Related event: Paper gives KV cache eviction a probabilistic foundation(3 posts)→

Original post →

More from Infra

Infra channel →