DeepSeek V4 training details: RL beyond collapse, WSD schedule, 1M context over 10T tokens

stochasticchasm · x · 2026-09-11

Community analysis suggests DeepSeek V4 is the second datapoint after MAI-thinking-1 of models continuing RL training beyond collapse — and, absurdly, merging checkpoints across different scaffolds/harnesses.

Follow-up notes: no instabilities observed, suggesting whatever issue arose during v4 training was solved — a good sign for the architecture; the team kept the WSD schedule and extended context length over 10T tokens, with training at 1M tokens for that long being quite wild. Unofficial analysis, unconfirmed by DeepSeek.

Related event: DeepSeek V4 training details: 10T+ tokens, 1M context, stable RL(3 posts)→

Original post →

More from Models

Models channel →