Estimating 800B Active Param Model: 10T Total Params if GPT-4 Style
_xjdr · x · 2026-07-24
Addressing rumors of an 800B active parameter model, developer @xjdr performed back-of-the-napkin calculations.
He suggests that if the model shares an architecture similar to the rumored GPT-4 rather than the DeepSeek-V3 sparse architecture, its total parameter count would be around 10 to 12 trillion. He notes that while unconfirmed, the source is typically well-informed and the math is sound.
Related event: 800B Active Parameter Model May Reach 10T Total(2 posts)→
More from Models
- Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model — teortaxesTex · 2026-09-11
- Pro 20x tier burns 60% of weekly quota in under a day with GPT-6 Astra — rschu · 2026-09-11
- Is DeepSeek's rumored K3 a scaled-down model, or something bigger? X users debate — teortaxesTex · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- 6TB of Fable data sold with leaked SSH keys, cloud creds tied to Xiaomi, Huawei, NIO — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11