4 sink tokens + 64-token window matches distilled linear attention, no training needed
burny_tech · x · 2026-09-14
The author reports that while millions have been spent post-training LLMs into linear attention to remove the KV cache bottleneck, a training-free trick—keeping just 4 initial tokens plus a 64-token sliding window—matches or beats distilled linear models on almost every benchmark.
The finding suggests much of the benefit of linearized models may come simply from preserving initial tokens (an attention-sink-like effect) rather than from the distillation itself, hinting at far cheaper alternatives to some KV-cache optimization work. Details in the thread.
More from Infra
- New research: standard SGD matches AdamW for LLM RL training, with far less memory overhead — zhaoran_wang · 2026-09-14
- OneLA Shares Linear-Attention States Across Beams for 1.54-2.46x Faster Generative-Rec Decoding — _reachsumit · 2026-09-14
- UBTech Opens World's First 10,000-Unit Humanoid Robot Factory; SK's AI Data Center Hits 900MW — 创业邦 · 2026-09-14
- Deep dive: thermal gradients are HBM's #1 scaling bottleneck, says new architecture breakdown — blaizedsouza · 2026-09-14
- Perplexity launches Hybrid Compute to split AI tasks between cloud and local Mac — Aiden_Tech_Ai · 2026-09-14
- RDNA4 local inference hits ~100 tok/s running Qwen3.8 Flash on dual R9700 — Public_Umpire_1099 · 2026-09-14