Rumor: Frontier Models May Be 48-Layer Transformers Looped Twice; DeepLoop Paper Explores Depth Scaling
dotey · x · 2026-09-06
- A circulating rumor (unverified) says some frontier models are essentially a 48-layer transformer looped twice (48L × 2); Chinese AI circles speculate OpenAI loops twice and ByteDance four times.
- The accompanying DeepLoop paper (Depth Scaling for Looped Transformers) presents methods to make looped transformers stable and scalable.
- If true, looped architectures would reshape assumptions about frontier parameter efficiency.
More from Models
- Leak: Grok Bot onboards ~100 customers as xAI sales enablement kicks off next week — Sauers_ · 2026-09-06
- "Total OpenAI victory" debate: coder says Cursor beats Codex harness by a wide margin — IndraVahan · 2026-09-06
- GLM Coding Plan ups Flash quotas: unlimited in ZCode, 2x elsewhere — pcuenq · 2026-09-06
- One-line verdict: astra fast ranks clearly above fable 5.1 — NERDDISCO · 2026-09-06
- Researchers dispute Anthropic's NEM reward hacking result as setup artifact — voooooogel · 2026-09-06
- Dev says OpenAI's Astra is first model making progress on his 'unreasonably complex' project — mrjonfinger · 2026-09-06