Would OpenAI bet $100M+ on Looped Transformer for Astra without scaling proof?

teortaxesTex · x · 2026-09-03

stalkermustang argues that if OpenAI's rumored >3T Astra model uses a Looped Transformer architecture, the company must have run extensive ablations and scaling studies before committing over $100M — so looping likely scales well on downstream metrics if not pretraining loss. A cited control experiment shows the tradeoff: a 355M Looped GPT (1 epoch, K=4) loses to a vanilla GPT trained 2 epochs at matched FLOPs.

Related event: Report: OpenAI's Astra Uses Recurrent Depth Architecture, Sparking Safety Concerns(11 posts)→

Original post →

More from Models

Models channel →