swallm Brings Sliding Window Attention to HF LLM Inference
A developer released swallm, an open-source project that adds sliding window attention to HuggingFace causal LLMs at inference time without retraining, cutting KV cache to about 3.5MB for 32K–64K contexts.
2026-09-06 ~ 2026-09-06 · 2 related posts