DeepSeek V4 training details: 10T+ tokens, 1M context, stable RL
Analyses reveal DeepSeek V4 trained on over 10T tokens with 1M context and no instability, using WSD scheduling, while k3's RL ran across 50 million sandboxes with millions concurrent—both continuing RL after collapse like MAI-thinking-1.
2026-09-11 ~ 2026-09-11 · 3 related posts
- Training run shows no instabilities and 1M context extended over ~10T tokens — stochasticchasm · 2026-09-11
- DeepSeek V4 training details: RL beyond collapse, WSD schedule, 1M context over 10T tokens — stochasticchasm · 2026-09-11
- k3 RL Run Reportedly Used ~50M Sandboxes With Millions Concurrent, Checkpoints Merged Across Scaffolds — stochasticchasm · 2026-09-11