Dev Ships Open-Source Sliding-Window Attention for HF LLMs, Hits 3.5MB KV Cache at 32K Context

ahsaor8 · reddit · 2026-09-06

The author turned their sliding-window attention (SWA) experiments into a reusable open-source project, swallm, for testing KV cache optimizations in long-context inference with HF causal LLMs.

Related event: swallm Brings Sliding Window Attention to HF LLM Inference(2 posts)→

Original post →

More from Infra

Infra channel →