Leak: DeepSeek V4.1 Pro ~2T params, trained sparse from scratch with no instabilities
teortaxesTex · x · 2026-10-03
Insider chatter (unverified): DeepSeek's V4.1 recipe reportedly shows no training instabilities and sparse-from-scratch works well, so they plan to scale it — V4.1 Pro is said to be 2T params (more with engrams), while the 8T model's design isn't settled. DeepSeek considers Kimi's approach inferior since Moonshot can't serve a 2.8T dense-compute model at volume, and Zhipu struggles with scaling; DS wants a recipe that scales yet stays cheap.
More from Infra
- GKE adds CPU Startup Boost: faster pod starts without over-provisioning — rseroter · 2026-10-03
- Huawei claims Ascend overtook Nvidia in China share; supply, not demand, is the bottleneck — teortaxesTex · 2026-10-03
- Turning an iPhone into a second GPU for a MacBook: 44% faster prefill on Qwen 27B — StayLameBro · 2026-10-03
- Micron CEO: memory supply will be much tighter in 2027-2028 than 2026 — dankvr · 2026-10-03
- mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking — StasBekman · 2026-10-03
- Measured on B200: nvfp4 is ~9% more efficient than mxfp4 with higher accuracy — pick nvfp4 on Blackwell — StasBekman · 2026-10-03