Schmidhuber: his 1991 fast weights work was the first Transformer variant, 30 years before GPT
SchmidhuberAI · x · 2026-09-19
- Responding to a claim that fly-style fast-weight continual learning is beyond current LLMs, Jürgen Schmidhuber points back to his 1991 Fast Weight Programmers research.
- The 1991 ULTRA model is mathematically equivalent to today's unnormalised linear Transformer: a feedforward net slowly learns to program the fast weights of another net, with compute scaling linearly rather than quadratically in input size.
- That same year he also introduced self-supervised pre-training for deep nets—the "P" in GPT. With compute a million times more expensive back then, the constraints produced a more efficient architecture, and the work ties into his broader metalearning agenda.
More from Research
- Google open-sources Fuse, a multi-agent framework for verifiable social reasoning in LLMs — google · 2026-09-19
- Study finds scientific beauty and impact are surprisingly correlated — jacobkimmel · 2026-09-19
- Trolley problem: Jev-style API turns jina-reranker-v3.5 into a ruthless decision engine — gaganghotra_ · 2026-09-19
- Probe guidance steers continuous diffusion LMs with just 1-3% extra inference compute — itsbautistam · 2026-09-19
- LLM Analysis of Wikipedia Battle Pages Crowns Napoleon the GOAT With 16.7 Career WAR — ctjlewis · 2026-09-19
- MIT paper tracks the birth, life and death of 108k open-weight AI models on Hugging Face — asusarla · 2026-09-19