Schmidhuber: Linearized self-attention Transformers date back to his 1992 NECO paper
SchmidhuberAI · x · 2026-10-05
Jürgen Schmidhuber re-upped his long-running priority claim on Transformers: his 1992 NECO paper already implemented Transformers with linearized self-attention, essentially equivalent to fast weight programmers (apart from normalization), with storage and control separated. Key/value was then called FROM/TO, and he says the attention terminology was introduced at ICANN 1993.
Related event: Schmidhuber Reiterates Linear-Attention Transformer Dates to 1991(3 posts)→
More from Research
- Shanghai AI Lab Proposes SCALE: Entropy-Gated Control to Reverse SFT Features — Shanghai-AI-Laboratory · 2026-10-05
- UCVG.cpp: Generate Control Vectors for Any LLM From a Single Prompt Pair — Egor4more · 2026-10-05
- How many digital minds on one GPU cluster? Synthese paper probes interwoven AI consciousness — burny_tech · 2026-10-05
- New preprint argues 'adaptive reframing'—revising problem representations—is a key unstudied dimension of intelligence — ValerioCapraro · 2026-10-05
- A Beautiful 2D Embedding Is Not Quantitative Evidence, Warns AI-for-Science Researcher — bravo_abad · 2026-10-05
- Stop Reading the Hessian as a Matrix: Eigenvalues Are Local Curvature of the Loss Surface — techNmak · 2026-10-05