RestoreKV: Recovering Performance Under Aggressive KV Cache Eviction
Changwoo Baek · hf · 2026-08-05
Aggressive KV cache eviction in long-context scenarios often leads to severe performance degradation. RestoreKV introduces a learned restoration mechanism to recover lost information without increasing the total KV budget.
- Core Mechanism: After context prefill, a few restore tokens attend to the full KV cache in a single LoRA-adapted pass to generate a compact, context-conditioned restore cache.
- Efficient Training: Trained via parameter-efficient self-distillation from a frozen full-cache model, optimizing only 0.4% of parameters with no task-specific tuning required.
- Results: On Qwen3-4B, it improves 59 of 60 paired settings. At a tight 5% budget, it boosts KVzip from 38.2 to 73.2 on RULER-4K, and achieves 86.4 accuracy on the KVPress Benchmark even at 16x compression.
More from Infra
- Anthropic to Develop Custom AI Chips for Faster and More Efficient Claude — Polymarket · 2026-08-05
- Wall Street Expects AI Capex Surge: OpenAI Quarterly Spending Could Top $18B — PTrubey · 2026-08-05
- Benchmarking MiniMax H3 on a 4090: Sage Attention Slashes Generation Time — thegr8anand · 2026-08-05
- 3090 Upgrade Dilemma: Is 24GB VRAM Enough or Jump to 32GB? — gtech02 · 2026-08-05
- Stanford Hazy Research: AI Agents Are Retiring CUDA Abstraction Layers — sumitdotml · 2026-08-05
- AI Agent Autonomously Writes Triton Kernel, Breaking NanoGPT Speedrun Record — cong_ml · 2026-08-05