A chronological map of looped-transformer papers, from Universal Transformers to DeepLoop

KyeGomezB · x · 2026-07-28

Chronological list of looped-transformer papers

The post collects a chronological reading list for Looped Transformers and related ideas, including:

The linked explainer argues that there are two broad ways to make an LLM “smarter”: scaling parameters or giving the model more compute/data during training. Looped transformer variants explore the second route by reusing layers or looping computation.

The image contrasts two patterns:

The cited Universal Transformers paper is highlighted as an early representative of this family, with claims about improved generalization on tasks that standard Transformers struggle with and the addition of dynamic halting.

Original post →

More from Research

Research channel →