Looped LMs could use 3x fewer parameters and a 3x smaller KV cache, matching standard Transformers

ChengleiSi · x · 2026-10-03

A research thread titled "Rethinking at Fixed Points" argues looped LMs don't have to pay for every loop: they can match a standard Transformer with 3x fewer parameters and a 3x smaller KV cache, scaling by adding FLOPs at constant memory.

The key is the fixed point of the loop. The authors introduce:

A long thread follows with full details.

Original post →

More from Research

Research channel →