Leaker estimates GPT-6 'Astra' as 4.2T-parameter looped language model
On September 28, leaker account scaling01 shared architectural speculation about OpenAI's next-generation model GPT-6 "Astra": a central estimate of roughly 4.2T total parameters, about 120B activated per token, and around 112 layers, with 50% of the layers looping once (looped language model). He said he won't publish his estimation method for now, but claimed evidence supports the speculation, and that the most common value in simulated inference runs was 4.5-4.6T.
Confirmed
- These are all personal estimates and speculation from scaling01, not official OpenAI information; he publicly stated the central estimates of 4.2T total parameters, 120B activation, about 112 layers, and 50% of layers looping once
- He added that if the total parameter count is much smaller (e.g., around 3T), each layer might need to loop 2-3 times; otherwise one would have to assume an activation ratio as high as 2%, which is hard to justify in practice
Not Yet Confirmed
- The model's actual parameter scale, layer count, and looping mechanism have not been verified by OpenAI, and the estimation method itself has not been disclosed
- Whether the model is a sparse MoE with many experts and a low activation rate is further speculation from the leaker
Why It Matters
- Looped architectures and sparse MoE are hot topics in current frontier model discussions; if GPT-6 truly adopts layer looping, it would mean trading fewer parameters for greater effective computational depth, offering useful insight into the scaling path of next-generation models
2026-09-28 ~ 2026-09-28 · 5 related posts
Primary sources
- [source] Leaker claims GPT-6 Astra is a 4.2T-param, 120B-active looped model; Opus 5.5 may not be looped — scaling01 · 2026-09-28
- [source] Leaker: GPT-6 at ~3T params would need 2-3x layer looping to stay plausible — scaling01 · 2026-09-28
3 near-duplicate retellings: scaling01 · scaling01 · scaling01