DeepSeek reportedly training a 2T-parameter model, with an 8T model on the roadmap
Terminator857 · reddit · 2026-09-21
Citing a circulating report, DeepSeek is training a 2T-parameter model and eventually plans an 8T-parameter model. For reference, DeepSeek Flash has 552B parameters and Pro has 1.6T total parameters with 49B activated weights per token; the rumored Mythos/Fable is estimated at 10T parameters. Unconfirmed by the company.
Related event: Rumor: DeepSeek Training 2T-Parameter Model with 8T on Roadmap(5 posts)→
More from Models
- Kev: open-source 0.8B/4B/9B judge models on Qwen3.5, the 9B fits a 32GB Mac — khiladi1729 · 2026-09-22
- Grok 4.7 still trails Muse, says user ranking top 3 as Fable, Astro, Muse — MicahBerkley · 2026-09-22
- OpenAI removes Ultrafast tier from GPT-5.6 Sol in Codex, fueling GPT-6 Sol rumors — imjustnewatai · 2026-09-22
- Mimo V2.6 undercuts Grok 4.7 by 6x on output price amid same-day model launches — op7418 · 2026-09-22
- Math community weighs in on AI 'Bel' claims: 100 solved problems, Millennium Problem skepticism — avaitopiper · 2026-09-22
- Terminal-Bench 4.0 leaderboard refresh draws attention to who's on top — ns123abc · 2026-09-22