NeurIPS 2025 paper: sparse attention emergence follows power laws, and repetition speeds it up
scychan_brains · x · 2026-09-02
A NeurIPS 2025 paper by Stephanie Chan and colleagues, "The emergence of sparse attention," studies how sparse attention patterns emerge over the course of Transformer training.
- Combining theoretical analysis of a toy model with empirical observations on small Transformers trained on a linear regression variant, the authors uncover the mechanics driving sparse attention emergence.
- Emergence timing follows power laws depending on task structure, architecture, and optimizer choice.
- Data repetition can greatly speed up emergence.
- Results are confirmed on an in-context associative recall task, offering a theoretically grounded framework for how data distributions and model design shape the learning dynamics behind one form of emergence.
More from Research
- Paper Shows On-Policy Distillation Gains Come from Self-Improvement, Not Teacher Guidance — _akhaliq · 2026-09-02
- The World Labs Introduces Atlas: Pixel-Perfect Camera Control World Model — Scobleizer · 2026-09-02
- 4DAnyone on Gradio: Optimized for <32GB VRAM, 8x Faster — _akhaliq · 2026-09-02
- Meta's Muse Voice Transcribe Balances Speed and Accuracy with Adaptive Delay — AIatMeta · 2026-09-02
- Speculative PTC: Overlapping tool calls with code generation for faster agents — a1zhang · 2026-09-02
- Replacing LLM heuristics with a Bayesian layer for PR defect prediction — Sakuraaa_29 · 2026-09-02