Schmidhuber: Linearized self-attention Transformers date back to his 1992 NECO paper

SchmidhuberAI · x · 2026-10-05

Jürgen Schmidhuber re-upped his long-running priority claim on Transformers: his 1992 NECO paper already implemented Transformers with linearized self-attention, essentially equivalent to fast weight programmers (apart from normalization), with storage and control separated. Key/value was then called FROM/TO, and he says the attention terminology was introduced at ICANN 1993.

Related event: Schmidhuber Reiterates Linear-Attention Transformer Dates to 1991(3 posts)→

Original post →

More from Research

Research channel →