Microsoft page reportedly confirms GPT-6 uses looped Transformers; LOOM paper scales to 9 loops
rickasaurus · x · 2026-10-07
A Microsoft public page is said to confirm OpenAI's GPT-6 series uses Looped Transformers (GPT-6.1 Sol runs 2 inference passes, with a mention of "instead of three"), seemingly validating The Information's earlier reporting. Companion arXiv paper LOOM (2610.01153) diagnoses why looped MoEs stall beyond two loops: amplified curse of depth destabilizing recurrence, and expert selection collapse where routers reuse the same experts. LOOM stabilizes recurrence via scaled residual updates and input re-injection, and diversifies computation with per-loop routers and a Looping Residual, sustaining improvements through 9 loops in FLOPs-matched experiments at 100M-1.7B scale.
Related event: Microsoft Page Reportedly Confirms GPT-6 Uses Looped Transformers(3 posts)→
More from Models
- Token-based pricing is strange: unpredictable, decoupled from value, misaligned incentives — amankhan · 2026-10-07
- JevBench splits leaderboard: open-weight models and API providers now ranked separately — airesearch12 · 2026-10-07
- Hot take: active params barely matter for cyber capability evals — RL env coverage is key — teortaxesTex · 2026-10-07
- Decagon launches Voice 3 with Chord voice model and duplex architecture for customer agents — Scobleizer · 2026-10-07
- JEV-9B, a Qwen3.5-based calibrated decision model, trends on Hugging Face — autotrust · 2026-10-07
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07