New Paper Derives a Learning Paradigm from the SGD–Linear Attention Duality

HanGuo97 · x · 2026-10-11

A new paper with Google PI exploits a dual view of gradient descent on linear layers: GD-trained layers compute exactly the same function as initial weights plus linear attention over their own training history. From this, the authors derive a new paradigm — instead of updating linear layers with GD each step, maintain a growing "KV cache" of gradient signals to attend over, letting the network grow during training.

Original post →

More from Research

Research channel →