Burkov: 2026 is putting recurrence back into the Transformer it removed in 2017

burkov · x · 2026-09-06

Andriy Burkov (author of The Hundred-Page Machine Learning Book) highlights an architectural full circle: the 2017 "Attention is All You Need" paper worked by removing recurrence from the then-SOTA LSTM-with-attention architecture, keeping only attention so parallelism enabled much bigger models. Now in 2026, the field is adding recurrence back into the Transformer to make models smarter without making them bigger.

Original post →

More from Research

Research channel →