Paper Lets Language Models Control Their Own Attention, Cutting Decode Cost 52%
KAIST's Declarative Attention lets language models declare which context regions matter during inference, skipping most KV cache reads and cutting decode cost by up to 52% in a zero-shot setting.
2026-09-03 ~ 2026-09-04 · 3 related posts
- KAIST's Declarative Attention lets LLMs skip most KV cache reads — kaist-ai · 2026-09-03
- Declarative Attention: LLMs Slash KV Cache Reads 31-52% Zero-Shot — chaumian · 2026-09-04
- Language models can control their own attention: 52% decoding cost cut on Gemma 4 31B — jm_alexia · 2026-09-04