Sliding-Window Attention at Inference Shrinks 64K-Context KV Cache to Just 3.5MB on Pretrained LLMs

ahsaor8 · reddit · 2026-09-06

The author implemented a sliding-window attention (SWA) inference layer for HuggingFace causal LLMs, with no model modification or retraining required.

Related event: swallm Brings Sliding Window Attention to HF LLM Inference(2 posts)→

Original post →

More from Infra

Infra channel →