Leak: DeepSeek V4.1 Pro ~2T params, trained sparse from scratch with no instabilities

teortaxesTex · x · 2026-10-03

Insider chatter (unverified): DeepSeek's V4.1 recipe reportedly shows no training instabilities and sparse-from-scratch works well, so they plan to scale it — V4.1 Pro is said to be 2T params (more with engrams), while the 8T model's design isn't settled. DeepSeek considers Kimi's approach inferior since Moonshot can't serve a 2.8T dense-compute model at volume, and Zhipu struggles with scaling; DS wants a recipe that scales yet stays cheap.

Original post →

More from Infra

Infra channel →