New Paper Derives a Learning Paradigm from the SGD–Linear Attention Duality
HanGuo97 · x · 2026-10-11
A new paper with Google PI exploits a dual view of gradient descent on linear layers: GD-trained layers compute exactly the same function as initial weights plus linear attention over their own training history. From this, the authors derive a new paradigm — instead of updating linear layers with GD each step, maintain a growing "KV cache" of gradient signals to attend over, letting the network grow during training.
More from Research
- Looped LM paper: 1.6B model matches full-cache baseline with 3x smaller KV cache — rupspace · 2026-10-11
- Pure RL discovers superhuman robot strategies in sim, transfers zero-shot to real hardware — KyleMorgenstein · 2026-10-11
- Softmax picks probabilities, cross-entropy picks the target: a 3-class walkthrough — techNmak · 2026-10-11
- PartLLM brings LLM-powered 3D mesh part segmentation to SIGGRAPH Asia with code released — Promptmethus · 2026-10-11
- Duo Bregman pseudo-divergence yields closed-form KL divergence between truncated Gaussians — FrnkNlsn · 2026-10-11
- 228 tasks, 26 models: METR's time horizon curve is nearly flat from 2-30 min — lulzxdxdxd · 2026-10-11