DeepLoop paper makes looped transformers scalable; rumor claims frontier models are 48 layers looped twice
StartupYou · x · 2026-09-03
yifanzhang claims 'some frontier models are basically a 48-layer transformer looped twice (48L x 2)'—unverified—and introduces DeepLoop: Depth Scaling for Looped Transformers, a paper making looped transformer training stable and scalable. If true, frontier labs may be reusing layers rather than merely stacking depth, but the rumor remains unconfirmed.
More from Models
- Anthropic staffer: new model feature live on API, coming to Claude Code within a day — trq212 · 2026-09-03
- 753B model thinks, 4B model writes: latent-space handoff matches frontier reasoning at 20x speed — TheMoonMidas · 2026-09-03
- Meta's Muse Spark 1.3 is now free on OpenCode — ramagetime · 2026-09-03
- Leaked: OpenAI's Secret 'Bel' Model Slated for Year-End After Astra's Agent-1 Stage — haider1 · 2026-09-03
- Reddit user on Gemini 3.8 Flash: more effort, similar result — Correct_Tomato1871 · 2026-09-03
- Qwen3.8-Flash-Next on 2x3090 + DDR4: expert cache PR lifts decode from 17 to 25-29 t/s — Extension-Bid-639 · 2026-09-03