New paper: tweaking recursive depth improves pre-training scaling exponents at scale

bo_wangbo · x · 2026-09-17

A new paper from Andrew G. Wilson's group shows that modifying recursive depth (looping) for model growth can improve scaling exponents in pre-training — compute-efficiency gains that increase with scale. It suggests architecture-level changes can beat simply pouring in more data and parameters.

Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→

Original post →

More from Research

Research channel →