Schmidhuber team asserts Linear Transformers replicate earlier Fast Weight Programmers

SchmidhuberAI · x · 2026-09-02

Jürgen Schmidhuber highlighted a citation dispute, referencing his team's 2021 paper 'Linear Transformers Are Secretly Fast Weight Programmers'. The work demonstrates the formal equivalence between linearized self-attention and 'Fast Weight Programmers' from the early 1990s. It also addresses memory capacity limitations in recent linear attention variants and proposes improved delta rule-like instructions.

Original post →

More from Research

Research channel →