Looped LMs at Fixed Points: 3x Smaller KV Cache, 1.79x Faster Prefill
IFM · hf · 2026-10-06
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
Every recurrence of a looped language model costs training, decoding, prefill, and RL. Key insight: the closer recurrent states get to fixed points, the less the path to them matters. This enables:
- Truncated backpropagation in training;
- Terminal KV sharing for decoding with almost no accuracy loss;
- A distilled student that prefills up to 1.79x faster;
- RL updates computed from saved rollout states, 2x faster than backpropagating through the trajectory.
The authors then improve the two components shaping fixed points:
- Depth prior: fixed-depth training breaks KV sharing; Huginn's broad prior dilutes supervision—instead, learn the prior from prediction feedback with an entropy term keeping it broad;
- Input injection: existing schemes let the state's input-aligned component amplify or cancel injection—orthogonal injection removes it.
Results: from 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale; at 1.6B, the learned prior with a 3x smaller KV cache matches fixed-depth training's downstream average with the full cache.
More from Models
- Daniel Han publishes summary of LLM benchmarks you can actually trust — danielhanchen · 2026-10-06
- Claim Verification Benchmarks Mostly Test Retrieval, Not Reasoning, Finds 24K-Trace Study — deliprao · 2026-10-06
- COLM26 study: LLMs ace claim verification benchmarks by taking shortcuts, not verifying — deliprao · 2026-10-06
- Opus 5.5 uses 26k tokens vs Astra's 12k yet costs 23% less per task at equal AA score — ChrisGPT · 2026-10-06
- GPT-6 Astra claimed to be first AI crossing world-class astrophysics threshold — johnseach · 2026-10-06
- $500/mo AI subscription is huge money in Jakarta: PPP pricing debate — sujingshen · 2026-10-06