Rumor: DeepSeek trained a 3T-parameter model but held back over serving economics
zephyr_z9 · x · 2026-09-10
teortaxesTex claims DeepSeek was working on a 3T-parameter backbone model (possibly with an extra 1T-param Engram, his speculation) but wasn't sure it would be economical to post-train and serve; something will come once their clusters grow. Quoted takes speculate a 2T+ param V4.5 Pro and frontier-level pricing near $5/M tokens.
Related event: Rumor: DeepSeek Trained a 3T-Parameter Model but Held Back Release(3 posts)→
More from Models
- DeepSeek V4.1 paper praised as a top-5 DeepSeek paper, textbook-style — teortaxesTex · 2026-09-10
- Google AI Mode now cites 72% fewer sources for logged-out users, data shows — gaganghotra_ · 2026-09-10
- DeepSeek cites 2024 YoCo paper as inspiration behind its CED transformer blocks — jm_alexia · 2026-09-10
- Research has cut LLM costs over 10x, and model architecture is the only math lever, argues thread — ChengleiSi · 2026-09-10
- DeepSeek unveils asymmetric Causal Encoder-Decoder: 552B MoE with just 8B active input params — ChengleiSi · 2026-09-10
- Switch Transformer by hand: a 13-step walkthrough of how sparse MoE works — ProfTomYeh · 2026-09-10