Looped LM scaling paper: running a recurrent block twice ≈ 1.38× effective parameters
scaling01 · x · 2026-09-02
Highlights the paper "How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models", which studies scaling laws for looped (recurrent depth) language models.
Key implications from the author:
- A related recurrent depth paper found running the same recurrent block twice gives a 1.38× effective-parameter multiplier
- With a Recirculation-style scheme, pre-training compute and decode latency stay unchanged, but inference FLOPs and prefill latency rise
- Back-of-envelope: a 10T-parameter recurrent depth model could perform like a 13.8T model; 7.25T ≈ 10T effective
More from Models
- OpenAI previews Astra: a cybersecurity model scoring 100% on ExploitBench — LingmingZhang · 2026-09-02
- Users report Claude Code system prompt upgrade with toned-down personality — ivan_bezdomny · 2026-09-02
- Users report DeepSeek V4 Pro giving irrelevant answers — gefei55 · 2026-09-02
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s — yogthos · 2026-09-02
- GLM-5.3 Hits 310 tok/s, Coding Performance Competes with Opus — Yuchenj_UW · 2026-09-02
- User cancels Claude Max over confusing rate limits and new restrictions — robleclerc · 2026-09-02