Training-free geometric KV routing cuts Qwen memory traffic by 16x
Electrical_Offer5667 · reddit · 2026-08-22
A training-free geometric KV routing method for frozen Qwen models assigns expired KV entries to centroid-represented regions, computing attention only over {sinks + recent window + selected regions}. Tests on Qwen3.5-2B (32k) and Qwen3-4B (8k) show 3.7x to 31x fewer physical KV reads while maintaining retrieval capabilities. Although wall-clock speed doesn't yet beat optimized dense attention due to GPU brute force and routing overhead, it highlights potential for reducing memory bandwidth and capacity in long-context scenarios. A GitHub demo is available.
More from Infra
- Nuclear 'hot rock' generates immense energy vs weak passive solar needing maintenance — tawnniee · 2026-08-22
- 2-4K GPUs can serve 100T tokens daily, sparking efficiency debate — teortaxesTex · 2026-08-22
- Berkeley/MIT Open Source Inference Engine: RTX 5090 Runs 284B Models at 25 tok/s — gnukeith · 2026-08-22
- Prediction market opens on whether Apple will announce 1TB+ unified memory chip — benfielding · 2026-08-22
- AI Value Shifts to Connectivity as HBM Density Peaks — BenBajarin · 2026-08-22
- DSPy 3.3.1 Released: Hardened Python Interpreter and MCP 2.0 Support — dbreunig · 2026-08-22