Burkov predicts looping recurrent 7B transformers will return and get good at coding

burkov · x · 2026-09-17

Andriy Burkov (author of The Hundred-Page Machine Learning Book) predicts that 7B models built on looping/recurrent transformer architectures with SOTA attention will soon return—and will actually be good for coding. It's a bet on architectural efficiency over scale for small coding models.

Original post →

More from AGI Musings

AGI Musings channel →