Looping Rewrite Scaling Exponents: 7.4B Model Matches GPT-3 13B With ~20x Less Compute
andrewgwils · x · 2026-09-17
A new paper from Andrew Gordon Wilson's group shows architectural interventions can modify pre-training scaling exponents: a 7.4B model-growth architecture via recursive depth (looping) matches GPT-3 13B on CORE with roughly 20x less compute, with efficiency gains that increase with scale. A simple boundary operator in vanilla transformers also helps, and looping regularizes in multi-epoch data-constrained settings.
Related event: Looped Depth Improves Scaling Exponents, 7.4B Model Matches GPT-3 13B(4 posts)→
More from Research
- Noam Brown on agent swarms, alignment, and recursive self-improvement — Recoil42 · 2026-09-18
- OpenAI is reportedly close to solving another Millennium Prize math problem — ResultBackground2450 · 2026-09-17
- STEER: steerable 3D head avatar motion prior for two-person conversation, code released — rsasaki0109 · 2026-09-17
- AI could power a global DNA sequencing network for real-time pathogen surveillance, argues ex-Google researcher — anshulkundaje · 2026-09-17
- Renormalizing probability scores destroys calibration, dev warns in model training debate — spikedoanz · 2026-09-17
- Noam Brown on agent swarms, recursive self-improvement, and alignment — Dwarkesh Podcast · 2026-09-17