Schmidhuber: Google's 2017 Transformer Builds on His 1991 Linear Attention Work
SchmidhuberAI · x · 2026-10-05
Schmidhuber reshares and stresses his 1991 "unnormalized linear Transformer" (ULTRA) work, arguing Google's 2017 normalized quadratic Transformer is based on its principles.
Key points:
- In the 1991 ULTRA, KEY/VALUE were called FROM/TO; its computational cost scales linearly with input size rather than quadratically.
- A 1993 paper on a recurrent ULTRA extension introduced the attention terminology: learning "internal spotlights of attention" by gradient descent.
- He points to the T in ChatGPT as descending from this early work, with links to details and references.
More from AGI Musings
- Dean Ball on the vulnerable world hypothesis: cognitive AI will unlock cheap ultra-destructive weapons — deanwball · 2026-10-05
- AI coding tools could be breaking the junior engineer pipeline — Suspicious_Orchid770 · 2026-10-05
- halvarflake: broken incentives since 2000 have ground down 2-3 generations of security engineers — halvarflake · 2026-10-05
- Before AI: 'Wikipedia Isn't a Valid Source.' After AI: 'Please Just Look at Wikipedia' — NC_Renic · 2026-10-05
- Dev builds AI skill cloning Matt Levine's writing style, won't release it over consent concerns — morqon · 2026-10-05
- Musk calls humanity a 'biological bootloader' for superintelligence, sparks backlash — wfithian · 2026-10-05