Looped Transformers: when is recurrent depth worth the extra compute?

ArchitectingAI · reddit · 2026-10-01

A study-group talk traces the path from Universal Transformers to looped Transformers and Huginn, where the same layers are reused for extra computation without added parameters. Key question: at a fixed parameter count, which tasks (iterative, compositional) justify the extra FLOPs and latency? Includes full slides and a call for compute-matched positive and negative results.

Original post →

More from Research

Research channel →