New Paper: Looping With Model Growth Bends Scaling Laws, Compute Gains Compound
akbirthko · x · 2026-09-17
The authors' new paper challenges the assumption that architectural changes only give constant-factor gains while pretraining progress comes mostly from data. They find that model growth, looping, and boundary operators yield compute multipliers over standard transformers that grow exponentially with each order of magnitude of compute:
- 1.55x at 1e20 FLOPs, projected 2.7x at 1e25
- Matches GPT-3 13B on CORE with 20x less compute
The core claim: the scaling exponent itself can be improved by architecture, with gains that compound with compute rather than staying constant.
Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→
More from Research
- Noam Brown on Dwarkesh: math progress hints at what happens when AI automates research — polynoamial · 2026-09-18
- ECCV attendees voice frustration: will foundation models absorb traditional CV research? — broodsugar · 2026-09-18
- Porting Agents' Last Exam's Linux CLI subset surfaced benchmark defects, yielding ALE-Gold — dejavucoder · 2026-09-18
- AISTATS 2027 opens submissions with AI review as a first-time feature — qberthet · 2026-09-18
- New claim: model architecture can improve scaling exponents — madhavsinghal_ · 2026-09-18
- Schmidhuber Revives 2020 RSI Talk, Says His Systems Learned Self-Improvement Since 1994 — SchmidhuberAI · 2026-09-18