KV Cache Scoring Is Useless: Random Eviction Matches Top Evictors at 32-43% Higher Throughput
burny_tech · x · 2026-09-04
A new paper, Random Attention, shows the dominant KV cache eviction paradigm—scoring cached tokens to decide what to keep—contributes almost nothing.
- Method: Keep the prompt's cache and evict reasoning tokens uniformly at random within each attention head, with no scoring at all
- Results: Matches the strongest prior evictor across 4 models and 6 reasoning tasks while delivering 32-43% higher throughput in vLLM deployment
- Why it works: The prompt is the fragile, critical part of the cache; reasoning traces protect themselves via textual redundancy and per-head copies, so random draws retain enough copies without any scoring
- Code and paper are publicly available
Related event: Study: Random KV Cache Eviction Beats Scoring Methods(2 posts)→
More from Infra
- Musk calls for a 'construction army' to staff xAI's Memphis supercomputer buildout — elonmusk · 2026-09-04
- Salesforce research: Random eviction of reasoning tokens matches selective KV cache compression — Salesforce · 2026-09-04
- Red Teamer Seeks Advice on Local LLM Harness for Authorized Security Exercises — CivilBug4007 · 2026-09-04
- Musk: winning AI means owning both chips and models — 'we must win here' — r0ck3t23 · 2026-09-04
- Gimlet Labs raises $300M at $3B valuation to split AI workloads across chip types — dinabass · 2026-09-04
- Kimi K3 Draft Collection: three open-sourced draft models trained with TorchSpec and vLLM on GB200 — AccBalanced · 2026-09-04