Recursive depth looping improves scaling exponents: 7.4B model matches GPT-3 13B with 20x less compute

andrewgwils · x · 2026-09-17

A new paper from Andrew Wilson's group shows that looping (recursive depth growth) during pre-training can improve the scaling exponent itself—not just constants—challenging the assumption that only data interventions move scaling exponents.

Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→

Original post →

More from Models

Models channel →