Meta's Loop Scaling Laws: Sparsity Gives ~3x Active-Param Efficiency, Recurrence ~2x on Reasoning
facebook · hf · 2026-10-01
Meta introduces Loop Scaling Laws, the first to jointly model recurrence and MoE sparsity alongside model size and data, via a bounded sparsity-conditional recurrence mapping. The laws predict looped-model loss better than prior alternatives and recover dense/MoE laws as special cases. Empirically: sparsity yields 3x active-parameter efficiency, recurrence 2x total-parameter efficiency on reasoning; at trillion-token scale, a law-derived looped MoE matches a 2x larger non-looped MoE at matched compute.
More from Infra
- Arduino argues sub-$900 embedded boards beat Mac minis for Physical AI agent economics — CatAstro_Piyush · 2026-10-01
- SemiAnalysis injects failures to rate GPU cluster renters — providers differ sharply on recovery — AccBalanced · 2026-10-01
- Redditor runs unattended DeepSeek loops for days: 237M tokens for just $3.48 — dogfoodarchitect · 2026-10-01
- Free course built from Cornell's GPU architecture workshop now shared publicly — idanbeck · 2026-10-01
- Qwen Flash Next MTP work resumes with official GGUF quants and llama.cpp PR — jacek2023 · 2026-10-01
- Undocumented Strata tip: set default sampling params via a sampling block in run config — KissMyShinyArse · 2026-10-01