Schmidhuber: 1991 ULTRA already had linear attention—the 'T' in ChatGPT traces to his 1991/1993 papers

SchmidhuberAI · x · 2026-10-05

Jürgen Schmidhuber reiterates that his 1991 unnormalized linear Transformer (ULTRA) featured linearized attention with linear rather than quadratic scaling in input size, predating the 2017 Transformer. He argues Google's 2017 normalized quadratic Transformer builds on ULTRA's principles (KEY/VALUE was then called FROM/TO), and that his 1993 recurrent extension coined the attention terminology.

Related event: Schmidhuber Reiterates Linear-Attention Transformer Dates to 1991(3 posts)→

Original post →

More from Research

Research channel →