A post estimates a rumored 800B-active model could imply 10T–12T total parameters
MoonL88537 · x · 2026-07-24
The reply says a rumored comment about an 800B-active model “checks out” mathematically and could imply a total size around 10T–12T parameters if it followed a GPT-4-like architecture instead of a sparse one.
The image quote argues that restraint, openness, and lower prices can actually increase the chance of making AGI, and even frames open source as part of that restraint strategy.
Related event: 800B Active Parameter Model May Reach 10T Total(2 posts)→
More from Models
- Teortaxes: Kimi stuck serving old K2 base as K2.8 rumored post-trained variant; DeepSeek's compute moat — teortaxesTex · 2026-09-11
- Qwen core dev teases 'v4p back? k3-0.2?' in cryptic model hint — JustinLin610 · 2026-09-11
- Frontier model weights are near-impossible to steal or self-replicate, argues Bindu Reddy vs AI doomers — bindureddy · 2026-09-11
- Why won't Google open source its STT models while startups ship SOTA open voice models? — techtotechbytechy · 2026-09-11
- User Praises DeepSeek's Model as Surprisingly Fast and Good in Hands-on Test — MaziyarPanahi · 2026-09-11
- Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model — teortaxesTex · 2026-09-11