DeepSeek V4 training details: RL beyond collapse, WSD schedule, 1M context over 10T tokens
stochasticchasm · x · 2026-09-11
Community analysis suggests DeepSeek V4 is the second datapoint after MAI-thinking-1 of models continuing RL training beyond collapse — and, absurdly, merging checkpoints across different scaffolds/harnesses.
Follow-up notes: no instabilities observed, suggesting whatever issue arose during v4 training was solved — a good sign for the architecture; the team kept the WSD schedule and extended context length over 10T tokens, with training at 1M tokens for that long being quite wild. Unofficial analysis, unconfirmed by DeepSeek.
Related event: DeepSeek V4 training details: 10T+ tokens, 1M context, stable RL(3 posts)→
More from Models
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11
- User switches back to GPT-5.6 Sol: barely uses quota and feels faster — CtrlAltDwayne · 2026-09-11
- Dev opinion: model differences shrink in a good harness; Grok 4.6 is good enough — gnukeith · 2026-09-11