Salesforce research: Random eviction of reasoning tokens matches selective KV cache compression
Salesforce · hf · 2026-09-04
Salesforce's "Random Attention" paper challenges conventional KV cache compression: with the prompt preserved, randomly evicting reasoning tokens performs on par with carefully scored selective compression methods.
- Key insight: reasoning traces are "self-protecting" — critical information is naturally redundant across long chains of thought, so random drops cause little unrecoverable loss
- This removes the need for per-token importance scoring, cutting overhead and simplifying cache management for long-CoT inference
- Directly relevant to serving-stack optimization for long-reasoning models: a much simpler eviction strategy achieves near-parity
Related event: Study: Random KV Cache Eviction Beats Scoring Methods(2 posts)→
More from Infra
- CXMT's 3D DRAM roadmap: risk production in 2028, mass production in 2029 — zephyr_z9 · 2026-09-04
- Musk calls for a 'construction army' to staff xAI's Memphis supercomputer buildout — elonmusk · 2026-09-04
- Red Teamer Seeks Advice on Local LLM Harness for Authorized Security Exercises — CivilBug4007 · 2026-09-04
- Musk: winning AI means owning both chips and models — 'we must win here' — r0ck3t23 · 2026-09-04
- Gimlet Labs raises $300M at $3B valuation to split AI workloads across chip types — dinabass · 2026-09-04
- Kimi K3 Draft Collection: three open-sourced draft models trained with TorchSpec and vLLM on GB200 — AccBalanced · 2026-09-04