Encoder-loop-decoder: a cleaner architecture proposed for looped transformers

letclaudiatweet · x · 2026-09-03

letclaudiatweet argues looping all transformer layers makes little sense: token-to-meaning and meaning-to-token mappings are deadweight after the first iteration. She proposes an encoder → looped middle section → decoder architecture, keeping representation bandwidth near the loop free. She admits knowing little of the looped-transformer literature and plans to read up.

Related event: Researchers Debate Looping Only Middle Layers in Recurrent Transformers(3 posts)→

Original post →

More from Research

Research channel →