Burkov: 2026 is putting recurrence back into the Transformer it removed in 2017
burkov · x · 2026-09-06
Andriy Burkov (author of The Hundred-Page Machine Learning Book) highlights an architectural full circle: the 2017 "Attention is All You Need" paper worked by removing recurrence from the then-SOTA LSTM-with-attention architecture, keeping only attention so parallelism enabled much bigger models. Now in 2026, the field is adding recurrence back into the Transformer to make models smarter without making them bigger.
More from Research
- NEAR AI's open-source Lean agent solves all of Putnam Bench for just $111 — lukaszkaiser · 2026-09-06
- KV Cache Explained: Why It's Crucial in LLM Inference and Often Misunderstood — techNmak · 2026-09-06
- PhD Student Uses Multi-Agent AI to Crack a 98-Year-Old Math Problem in 48 Hours — 量子位 · 2026-09-06
- New piece: Cognitive maps as a medium for thought — abenitezburraco · 2026-09-06
- Google paper mathematically shows test-time compute backfires when training data lacks the skill — solyarisoftware · 2026-09-06
- MAVIN: multi-shot audio-video generation with narrative control, ECCV 2026 Oral — jiqizhixin · 2026-09-06