Schmidhuber: His 1991 Paper Introduced Pre-training, Positional Encoding and Distillation

SchmidhuberAI · x · 2026-10-06

Jürgen Schmidhuber restated his long-running priority claim that his 1991 TUM technical report Neural sequence chunkers (journalized in Neural Computation, 1992) introduced three techniques now essential to modern LLMs: pre-training for deep neural nets (the "P" in ChatGPT), positional encoding (Sec. 5.1), and knowledge distillation in neural nets (Sec. 3.2.2 & 4), with links to the original report, journal paper, and overview.

Original post →

More from Research

Research channel →