Looped LMs at Fixed Points: 3x Smaller KV Cache, 1.79x Faster Prefill

IFM · hf · 2026-10-06

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Every recurrence of a looped language model costs training, decoding, prefill, and RL. Key insight: the closer recurrent states get to fixed points, the less the path to them matters. This enables:

The authors then improve the two components shaping fixed points:

Results: from 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale; at 1.6B, the learned prior with a 3x smaller KV cache matches fixed-depth training's downstream average with the full cache.

Original post →

More from Models

Models channel →