Declarative Attention: LLMs Slash KV Cache Reads 31-52% Zero-Shot
chaumian · x · 2026-09-04
The arXiv paper "Language Models Can Control Their Own Attention" introduces Declarative Attention (DA), a protocol where the model itself declares where to attend in its chain-of-thought via <global>/<focus>/<local> modes, letting the inference engine skip most KV cache reads. Zero-shot across 15 long-context tasks, DA cuts total attended tokens during decoding by 52.0% (Gemma-4-31B) and 31.1% (Qwen-3.6-27B) with accuracy drops of only 1.27pp/2.75pp that shrink with scale — a new axis of sparse attention.
More from Research
- Google DeepMind Launches WeatherNext 3, Its Most Advanced Global Weather AI Model — rseroter · 2026-09-04
- New Model Release Features Quantum-Inspired Weight Permutation Technique — CamachoCollados · 2026-09-04
- Hobbyist says nightly RL on just 400-600 rollouts makes model gains noticeably deployable — cephaloform · 2026-09-04
- Baseten launches Base Labs, an open-source AI research lab publishing everything including failures — eigenron · 2026-09-04
- Researcher: All model and optimizer hyperparameters are functions of width and tokens — dlwh · 2026-09-04
- OpenAI quietly publishes Lean proof repos, seen as warm-up for Astra release — NoFaithlessness951 · 2026-09-04