Declarative Attention Cuts LLM Attention Work by 52%

rohanpaul_ai · x · 2026-09-05

A Google deployment paper introduces Declarative Attention: instead of rereading the whole context for every token, the model itself declares which context regions it needs and the inference engine skips the rest—no external scorer required. Zero-shot on 15 long-context tasks, attended tokens fell 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B, with accuracy drops of just 1.27pp and 2.75pp. Larger models handle the trade-off better, suggesting further gains with training.

Related event: Google paper proposes Declarative Attention, cutting attention cost by 52%(2 posts)→

Original post →

More from Infra

Infra channel →