New paper: recursive looping boosts pre-training scaling exponents

burny_tech · x · 2026-09-20

Andrew Wu (@andrewgwils) and collaborators (@charllechen, @akshayvegesna, @industriaalist) released a new paper showing that modifying recursive depth (looping) for model growth can improve scaling exponents in pre-training, yielding compute efficiency gains that grow with scale. In a quoted thread, @awesomeruler notes that scaling the simple recipe from their slowrun record (growing loops + input injection) with untied weights beats vanilla models' scaling exponent, especially under data-constrained settings.

Original post →

More from Research

Research channel →