ExANS: Open-Source Lossless KV Cache Compression Hits 622 GB/s on H100
arnav__1 · hn · 2026-08-06
OpenLake has introduced ExANS, an open-source GPU lossless compression scheme for LLM KV caches.
- Core Insight: While BF16 tensors are typically hard to compress due to high entropy, real-world KV blocks have very low entropy in the exponent byte. The solution targets only the exponent stream for compression.
- Performance: Tested on an H100 with production KV snapshots, it achieves 1.51x lossless compression and a median GPU decode speed of 622 GB/s.
- Engineering Value: Decompression is 10x faster than a 400 Gbps NIC bandwidth, effectively alleviating PCIe/NIC bottlenecks and reducing Time To First Token (TTFT) without quality degradation.
- Integration: Available in OpenLake v0.8 with vLLM and SGLang connectors, requiring zero changes in the inference engine.
More from Infra
- Jeff Dean and Scientists Left Google Citing TPU Infrastructure Limits on Research — firstadopter · 2026-08-06
- ARCHead: New LLM Output Head Quantization Method Substantially Reduces Storage with Minimal Loss — Şuayp Talha Kocabay · 2026-08-06
- Can 8x NVIDIA V100 GPUs Handle DeepSeek Inference for a 50-Person Team? — MKU64 · 2026-08-06
- Is Upgrading to 96GB RAM Worthwhile for RTX 5090 Local AI Workflows? — Beastly4k · 2026-08-06
- NVIDIA Unveils RTX Spark Superchip: 1 Petaflop of FP4 AI Power for Next-Gen PCs — nvidia · 2026-08-06
- New ComfyUI Nodes Boost Minimax H3 4x, Krea 2 5.6x Faster — Certain-Will-2769 · 2026-08-06