Would OpenAI bet $100M+ on Looped Transformer for Astra without scaling proof?
teortaxesTex · x · 2026-09-03
stalkermustang argues that if OpenAI's rumored >3T Astra model uses a Looped Transformer architecture, the company must have run extensive ablations and scaling studies before committing over $100M — so looping likely scales well on downstream metrics if not pretraining loss. A cited control experiment shows the tradeoff: a 355M Looped GPT (1 epoch, K=4) loses to a vanilla GPT trained 2 epochs at matched FLOPs.
More from Models
- Tencent Hunyuan 770B compressed from ~1.5TB to ~214GiB with mixed quantization — Aiden_Tech_Ai · 2026-09-03
- Flappy Bird coding test: Kimi K3 costs 9x more than DeepSeek V4 Flash — Arindam_1729 · 2026-09-03
- 33.3% of ImageNet images contain more than one class — the answer key itself is flawed — RexDouglass · 2026-09-03
- Cohere Labs head: multilingual AI's 'longitude problem' — scale isn't enough — Cohere_Labs · 2026-09-03
- Gemini keeps generating images even when users explicitly say stop — a tool-calling UX bug report — JacketMajor6561 · 2026-09-03
- antirez: Zero Out Steering When Possible—It Always Adds Distortion — antirez · 2026-09-03