Training run shows no instabilities and 1M context extended over ~10T tokens
stochasticchasm · x · 2026-09-11
- No training instabilities observed so far, suggesting the issue from v4 training has been resolved — a good sign for the architecture.
- Notably, the team continues with the WSD learning rate schedule and extends context length over a 10T token training run.
- The author finds it remarkable that training is sustained at 1M token context for that long.
Related event: DeepSeek V4 training details: 10T+ tokens, 1M context, stable RL(3 posts)→
More from Infra
- Google Commits $15B to AI Infrastructure Buildout in Finland — LinkedInNews · 2026-09-11
- DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression — ollama · 2026-09-11
- DIY-friendly KiCad footprints for AI MELF resistors, milled at home — debreuil · 2026-09-11
- Inference providers barely break even: $10K revenue yields just $200 profit — metalvendetta · 2026-09-11
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11