Survey of Attention Evolution: Contextual Memory Becomes the Core of LLM Architecture Design
Zhentao Tan · hf · 2026-10-01
A survey tracing the evolution of attention in LLMs:
- Self-attention enables fine-grained, query-dependent context access but incurs quadratic prefill cost and a growing KV cache, spurring research on explicit-memory compression, sparse access, recurrent state, structured state dynamics, and heterogeneous mechanism composition.
- It proposes a five-dimensional lens — Memory Representation, Update, Access, Readout, Integration — to compare overlapping research lines.
- Reconstructing developments from 59 release-level records across 14 model lineages plus 11 open-weight endpoints, it finds explicit-memory and recurrent-state methods converging on overlapping memory functions, and heterogeneous architectures increasingly coordinating memory processing across network depth.
- It advances a stateful multidimensional memory-routing hypothesis: persistent memory is organized by temporal scope, depth, substrate, and granularity, with coordinated Sparse Write/Read governing retention and contribution. Efficient sequence architecture is shifting from optimizing a single attention operator to organizing contextual memory.
More from Research
- LANTERN uses LLM internal activations to surface four novel OEIS integer sequence relations in under 8 hours — Pavel Tikhonov · 2026-10-01
- The geometry of inference in transformer residual streams: how predictions sharpen with depth — Timur Mudarisov · 2026-10-01
- Kernelised functional Bregman divergences paper accepted at NeurIPS 2026 NeurReps workshop — FrnkNlsn · 2026-10-01
- Why vision and hearing dominate experience: bigger cortical regions, not a saliency system — Sauers_ · 2026-10-01
- Hidden Dates in System Prompts Swing LLM Eval Scores by Up to 14% — Mario Sanz-Guerrero · 2026-10-01
- CheatBench Launches to Measure Reward Gaming and Cheating in AI Agents — cais · 2026-10-01