New Paper: Looping With Model Growth Bends Scaling Laws, Compute Gains Compound

akbirthko · x · 2026-09-17

The authors' new paper challenges the assumption that architectural changes only give constant-factor gains while pretraining progress comes mostly from data. They find that model growth, looping, and boundary operators yield compute multipliers over standard transformers that grow exponentially with each order of magnitude of compute:

The core claim: the scaling exponent itself can be improved by architecture, with gains that compound with compute rather than staying constant.

Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→

Original post →

More from Research

Research channel →