Declarative Attention lets LLMs declare their own focus, cutting 52% of KV cache reads

eigenlaplace · reddit · 2026-09-05

A new arXiv paper proposes Declarative Attention (DA): instead of pre-selecting tokens with extrinsic proxy scores, the model itself declares where it needs to attend within its chain-of-thought.

The authors frame DA as a new axis of sparse attention with further potential under training-based methods.

Related event: Declarative Attention Lets LLMs Cut KV Reads by 52%(3 posts)→

Original post →

More from Infra

Infra channel →