DeepSeek V4 training details: 10T+ tokens, 1M context, stable RL

Analyses reveal DeepSeek V4 trained on over 10T tokens with 1M context and no instability, using WSD scheduling, while k3's RL ran across 50 million sandboxes with millions concurrent—both continuing RL after collapse like MAI-thinking-1.

2026-09-11 ~ 2026-09-11 · 3 related posts