Looping Rewrite Scaling Exponents: 7.4B Model Matches GPT-3 13B With ~20x Less Compute

andrewgwils · x · 2026-09-17

A new paper from Andrew Gordon Wilson's group shows architectural interventions can modify pre-training scaling exponents: a 7.4B model-growth architecture via recursive depth (looping) matches GPT-3 13B on CORE with roughly 20x less compute, with efficiency gains that increase with scale. A simple boundary operator in vanilla transformers also helps, and looping regularizes in multi-epoch data-constrained settings.

Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→

Original post →

More from Research

Research channel →