Schmidhuber: His 1991 Paper Introduced Pre-training, Positional Encoding and Distillation
SchmidhuberAI · x · 2026-10-06
Jürgen Schmidhuber restated his long-running priority claim that his 1991 TUM technical report Neural sequence chunkers (journalized in Neural Computation, 1992) introduced three techniques now essential to modern LLMs: pre-training for deep neural nets (the "P" in ChatGPT), positional encoding (Sec. 5.1), and knowledge distillation in neural nets (Sec. 3.2.2 & 4), with links to the original report, journal paper, and overview.
More from Research
- Ben Goertzel's d-calculus: the math of goal preservation under AI self-improvement — burny_tech · 2026-10-06
- Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason — arankomatsuzaki · 2026-10-06
- Schmidhuber explains his 1991 positional encoding: hyperbolic 1/t time-decay still widely used — SchmidhuberAI · 2026-10-06
- GDELT uses Gemini 3 to reason over 25 years of TV news across 75 countries and 150 languages daily — rseroter · 2026-10-06
- Researcher flags that new APSP construction doesn't beat trivial quantum n^2.5 bound — aran_nayebi · 2026-10-06
- Meta AI's MIRA splits research agents to fix long-horizon decision learning — dair_ai · 2026-10-06