50M-token persistent KV memory on one H100: 2.8-4.3x faster, 8.8-12.3x less GPU energy
Corbenci · hf · 2026-10-09
A Hugging Face report tests galahad-kv, an open package that saves 16K-token KV blocks to encrypted local NVMe and reloads them byte-exact instead of recomputing. On 50M tokens of real text served via vLLM on a single H100 (Gemma 4 12B/31B), all 100 probed blocks loaded with zero recompute; loading was 2.8-4.3x faster than recompute, used 8.8-12.3x less GPU energy, and kept GPU memory flat. Planted-fact recall hit 82/100 (12B) and 98/100 (31B) with no hallucinations. Limits: it's state reuse, not a wider attention window; needs TBs of NVMe. Includes a gaming-resistant test protocol and single-GPU reproduction.
More from Infra
- Self-hosted SearXNG behind 67 proxy chains, accessible only via tailnet — haydendevs · 2026-10-09
- iPad + AMD R9700 eGPU Runs Qwen3.8-27B at 159 tok/s via Open-Source LSE Engine — TheOriginalG2 · 2026-10-09
- GlobalFoundries Signs 5-Year TSMC Deal to Make AI Chip Interposers in New York — teortaxesTex · 2026-10-09
- NVIDIA pitches RTX Spark PCs for local fine-tuning, AI coding agents and deployment — danielhanchen · 2026-10-09
- Air Street GP: failed frontier labs are pivoting to data centers serving Chinese models — johncoogan · 2026-10-09
- Domingos: China zooms past America in data center capacity — pmddomingos · 2026-10-09