Looped LMs could use 3x fewer parameters and a 3x smaller KV cache, matching standard Transformers
ChengleiSi · x · 2026-10-03
A research thread titled "Rethinking at Fixed Points" argues looped LMs don't have to pay for every loop: they can match a standard Transformer with 3x fewer parameters and a 3x smaller KV cache, scaling by adding FLOPs at constant memory.
The key is the fixed point of the loop. The authors introduce:
- Fixed-Point Shortcuts: what fixed points buy a looped model in training (pre-training, post-training) and inference (prefill, decoding)
- A learned depth prior and orthogonal injection: two ways to shape those fixed points
A long thread follows with full details.
More from Research
- Unitree's UnifoLM-WLA-1.0: one 6B model for 64 whole-body humanoid tasks — WebAssemblyMan · 2026-10-03
- Mathematician Kontorovich admits he was wrong about AI autonomously formalizing math — AlexKontorovich · 2026-10-03
- New paper: training LLMs to verbalize when they know they're being evaluated — xuanalogue · 2026-10-03
- Schmidhuber Replies to the Pope With His 2008 Curiosity Paper on Compression Progress — prajdabre · 2026-10-03
- How DatologyAI Generated 12 Trillion Synthetic Tokens — And Fixed 4 Pipeline Bottlenecks — AI Engineer · 2026-10-03
- CMU team releases SMDD-Bench: 502 small-molecule drug design tasks for RL agents — willcb · 2026-10-03