Declarative Attention Cuts LLM Attention Work by 52%
rohanpaul_ai · x · 2026-09-05
A Google deployment paper introduces Declarative Attention: instead of rereading the whole context for every token, the model itself declares which context regions it needs and the inference engine skips the rest—no external scorer required. Zero-shot on 15 long-context tasks, attended tokens fell 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, with accuracy drops of just 1.27pp and 2.75pp. Larger models handle the trade-off better, suggesting further gains with training.
Related event: Google paper proposes Declarative Attention, cutting attention cost by 52%(2 posts)→
More from Infra
- Hybrid Compute on Mac ships with open-sourced local inference engine and PII classifier — andrewgwils · 2026-09-05
- Tesla's RIM process kills the paint shop, shrinking Cybercab factory footprint ~50% — elonmusk · 2026-09-05
- AMD unveils Threadripper Halo Station: 96-core CPU plus MI350P cards with 288GB HBM3E — Aroochacha · 2026-09-05
- Extropic unveils Z1T, first transformer family for sparse probabilistic hardware with up to 140x GPU energy efficiency — afurgs · 2026-09-05
- Why would Google partner with independent AI web indexes like Parallel? — Genzinvestor16180339 · 2026-09-05
- PwC: every $1 of data center construction commits ~$12 to future ICT equipment — rohanpaul_ai · 2026-09-05