Sliding-Window Attention at Inference Shrinks 64K-Context KV Cache to Just 3.5MB on Pretrained LLMs
ahsaor8 · reddit · 2026-09-06
The author implemented a sliding-window attention (SWA) inference layer for HuggingFace causal LLMs, with no model modification or retraining required.
- Core idea: The KV cache keeps only attention sinks plus a fixed-size sliding window, stored in a ring buffer, combined with streaming prefill and chunked attention masks.
- Benchmarks (Qwen2.5-7B): At 16K context, KV memory drops from 923MB to 3.5MB; at 32K, from 1.84GB to 3.5MB; at 64K, full attention OOMs while SWA stays bounded. At 16K, TPOT falls from 38.4ms to 30.5ms.
- Trade-off: Retrieval of information far outside the window degrades; the author is investigating how much is inherent to SWA versus an artifact of the implementation or model.
- The code is open source, and the author is asking for feedback: which architectures to validate next, which failure cases to test, and how to integrate with existing HF inference pipelines.
Related event: swallm Brings Sliding Window Attention to HF LLM Inference(2 posts)→
More from Infra
- Lightpanda: Open-Source Headless Browser for AI Agents, 11x Faster Than Chrome — Shruti_0810 · 2026-09-06
- Baseten's Philip Kiely Launches Inference Engineering Book, Plus Learning Resources — kmeanskaran · 2026-09-06
- Polygres turns your Postgres into a hybrid search context layer for AI agents — Scobleizer · 2026-09-06
- TCS may invest up to $7.4 billion with TPG in a gigawatt AI campus in Hyderabad — emmanuelvivier · 2026-09-06
- KV cache pressure tool shows vLLM's advertised 2M-token cache can retain 3M after fixes — t4a8945 · 2026-09-06
- Independent Sweep Puts DiffusionGemma Peak Throughput at 3k tok/s, 3x Paper's Claim — bodonoghue85 · 2026-09-06