VaSE: training-free stochastic KV cache eviction for reasoning models, shown at COLM 2026
robinomial · x · 2026-10-05
Presented at the Efficient Reasoning Workshop at COLM 2026 (Oct 9, 12:30–13:30), VaSE (Value-Aware Stochastic KV Cache Eviction) tackles KV cache bloat from long chain-of-thought reasoning. Eviction typically caps memory but hurts capability; this training-free recipe keeps large-magnitude value states and evicts stochastically, cutting memory cost with less capability drop. Code is available. The author is also seeking Spring/Summer 2027 internships.
More from Infra
- Inference Economics: Google Now Processes 3.2 Quadrillion Tokens Monthly, 7x a Year Ago — mikeflache · 2026-10-05
- InP choke pushes industry toward 1um lasers on GaAs and hollow-core fiber — jwt0625 · 2026-10-05
- Cloudflare's new Web Search API is just a wrapper around Exa and a few other providers — gaganghotra_ · 2026-10-05
- Top 10 providers on OpenRouter ranked by monthly token volume — stuffyokodraws · 2026-10-05
- Crusoe CEO explains its layered compute business: 5-year leases to high-margin inference — AccBalanced · 2026-10-05
- AMD MI355X beats Nvidia on inference margins in SemiAnalysis InferenceX benchmark — AccBalanced · 2026-10-05