Meta's Loop Scaling Laws: Sparsity Gives ~3x Active-Param Efficiency, Recurrence ~2x on Reasoning

facebook · hf · 2026-10-01

Meta introduces Loop Scaling Laws, the first to jointly model recurrence and MoE sparsity alongside model size and data, via a bounded sparsity-conditional recurrence mapping. The laws predict looped-model loss better than prior alternatives and recover dense/MoE laws as special cases. Empirically: sparsity yields 3x active-parameter efficiency, recurrence 2x total-parameter efficiency on reasoning; at trillion-token scale, a law-derived looped MoE matches a 2x larger non-looped MoE at matched compute.

Original post →

More from Infra

Infra channel →