DeepSeek reportedly training a 2T-parameter model, with an 8T model on the roadmap

Terminator857 · reddit · 2026-09-21

Citing a circulating report, DeepSeek is training a 2T-parameter model and eventually plans an 8T-parameter model. For reference, DeepSeek Flash has 552B parameters and Pro has 1.6T total parameters with 49B activated weights per token; the rumored Mythos/Fable is estimated at 10T parameters. Unconfirmed by the company.

Related event: Rumor: DeepSeek Training 2T-Parameter Model with 8T on Roadmap(5 posts)→

Original post →

More from Models

Models channel →