50M-token persistent KV memory on one H100: 2.8-4.3x faster, 8.8-12.3x less GPU energy

Corbenci · hf · 2026-10-09

A Hugging Face report tests galahad-kv, an open package that saves 16K-token KV blocks to encrypted local NVMe and reloads them byte-exact instead of recomputing. On 50M tokens of real text served via vLLM on a single H100 (Gemma 4 12B/31B), all 100 probed blocks loaded with zero recompute; loading was 2.8-4.3x faster than recompute, used 8.8-12.3x less GPU energy, and kept GPU memory flat. Planted-fact recall hit 82/100 (12B) and 98/100 (31B) with no hallucinations. Limits: it's state reuse, not a wider attention window; needs TBs of NVMe. Includes a gaming-resistant test protocol and single-GPU reproduction.

Original post →

More from Infra

Infra channel →