DeepMind's Declarative Attention: model declares which KV cache regions to read, skipping most reads
omarsar0 · x · 2026-09-04
- Google DeepMind and colleagues propose Declarative Attention: instead of proxy scores guessing relevant tokens at O(N) per step, the model declares within its chain-of-thought where it needs to look.
- The inference engine parses these declarations like tool calls and skips most KV cache reads, splitting generation into global, focus, and local modes.
- Zero-shot on off-the-shelf weights across 15 long-context tasks, attended tokens drop dramatically.
Related event: Declarative Attention Lets LLMs Control Their Own KV Cache Reading(5 posts)→
More from Infra
- Four Mac Studios Over Thunderbolt Run Trillion-Parameter Model on One Wall Outlet as Apple Pitches Local AI — mark_k · 2026-09-23
- Archgen Labs, aiming to make chip design 1,000x faster, lands YC backing — retr0jirachi · 2026-09-23
- Running Omarchy and a 90M-parameter LLM on a PSP — Kyrannio · 2026-09-23
- A Redditor built a dual AMD R9700 local inference rig and crowdsources tuning advice — Current-Ticket4214 · 2026-09-23
- OpenRouter Batch API spans 71 models, auto-routes to cheapest provider — jeff_weinstein · 2026-09-23
- Fully Offline NotebookLM Alternative: Ollama + Open WebUI + RAG Stack Suggested — betobagio · 2026-09-23