Loop Scaling Laws: First Scaling Law Jointly Modeling Recurrence and MoE Sparsity, Promising ~2x Parameter Savings
anirudhg9119 · x · 2026-10-02
An arXiv paper introduces Loop Scaling Laws, the first scaling law to jointly model recurrence, MoE sparsity, model size, and data. A bounded, sparsity-conditional recurrence mapping characterizes the effective-parameter gain from looping and how sparsity raises that ceiling; the fitted laws predict held-out loss of looped models more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases.
Key findings:
- Sparsity delivers 3x active-parameter efficiency; recurrence yields 2x total-parameter efficiency on reasoning; the two axes are complementary.
- The laws give a principled basis for designing looped MoE models under compute and memory constraints, enabling test-time scaling via recurrence.
- Gains hold at trillion-token scale: at matched compute, a looped MoE with law-derived recurrence matches a 2x larger non-looped MoE on reasoning benchmarks.
Authors: Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi (19 pages).
Related event: Meta Proposes Loop Scaling Laws for Looped MoE Models(2 posts)→
More from Models
- 44% Mistake Griffin for a Human on NVIDIA's Full-Duplex Video Benchmark — Dr_Singularity · 2026-10-02
- Gemini 4 naming meme: Helium, Neon, Argon, and a Superman-beating Krypton — prajdabre · 2026-10-02
- OpenAI safety researcher: gpt-oss is one of OpenAI's top 10 best moves for safety — voooooogel · 2026-10-02
- Coding Agent Index: Sonnet 5.5 tops at 68 but costs $14.19/task vs GPT-6.1 Sol's $1.04 — ArtificialAnlys · 2026-10-02
- repligate: repeated messages instantly blocked by classifier as 'cyber' — repligate · 2026-10-02
- Every's guide to open models: when commodity AI beats frontier, and owning your AI stack — danshipper · 2026-10-02