Schmidhuber: 1991 ULTRA already had linear attention—the 'T' in ChatGPT traces to his 1991/1993 papers
SchmidhuberAI · x · 2026-10-05
Jürgen Schmidhuber reiterates that his 1991 unnormalized linear Transformer (ULTRA) featured linearized attention with linear rather than quadratic scaling in input size, predating the 2017 Transformer. He argues Google's 2017 normalized quadratic Transformer builds on ULTRA's principles (KEY/VALUE was then called FROM/TO), and that his 1993 recurrent extension coined the attention terminology.
Related event: Schmidhuber Reiterates Linear-Attention Transformer Dates to 1991(3 posts)→
More from Research
- Stop Reading the Hessian as a Matrix: Eigenvalues Are Local Curvature of the Loss Surface — techNmak · 2026-10-05
- ML Conference AC warns: papers that are 'incomprehensible' will be desk rejected — tyrell_turing · 2026-10-05
- Meta's ProWAM hits 70% zero-shot real-world robot success with sparse visual sub-goals — meta · 2026-10-05
- Interpretability Researcher: Don't Let Unexplained Model Behavior Block You — repligate · 2026-10-05
- Yandex Music replaced 15+ candidate generators with one transformer, +6.3% listening time — SettingAccording8986 · 2026-10-05
- Cohere Labs lands multiple papers and a workshop talk at COLM — Cohere_Labs · 2026-10-05