SeKV: Adaptive KV Cache for Long-Context Inference

microsoft · hf · 2026-07-08

Microsoft proposes SeKV, a resolution-adaptive semantic KV cache method. It compresses context into entropy-guided spans with hierarchical storage across GPU-CPU memory. This enables efficient long-context processing with minimal memory overhead while preserving token-level details.

Original post →

More from Infra

Infra channel →