Schmidhuber team asserts Linear Transformers replicate earlier Fast Weight Programmers
SchmidhuberAI · x · 2026-09-02
Jürgen Schmidhuber highlighted a citation dispute, referencing his team's 2021 paper 'Linear Transformers Are Secretly Fast Weight Programmers'. The work demonstrates the formal equivalence between linearized self-attention and 'Fast Weight Programmers' from the early 1990s. It also addresses memory capacity limitations in recent linear attention variants and proposes improved delta rule-like instructions.
More from Research
- Analysis of ExploitGym: OpenAI Model Used Specific Vulnerabilities for Hacking — BlackHC · 2026-09-02
- Lapis (ECCV 2026) Open Sourced: Pixel-Space Diffusion Depth Estimation with Linear Attention — kwangmoo_yi · 2026-09-02
- Lapis: Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention — kwangmoo_yi · 2026-09-02
- MiniMax Releases H3-World: An Interactive World Model — CryptoBeth96 · 2026-09-02
- Open Source DCLAP Model and SAE Analysis Tool to Fix Long-Tail Terms in Music Search — Old_Rock_9457 · 2026-09-02
- Dev notes record outcomes, but the reasoning dies in the transcript — Sea-Perception1619 · 2026-09-02