DeepSeek-V4 Writeup Reveals Use of Agentic Trajectories in Mid-Training
cwolferesearch · x · 2026-08-06
A researcher discovered upon re-reading the DeepSeek-V4 writeup that the team incorporated a certain ratio of agentic trajectories into the mid-training phase.
Notably, this is the only sentence in the entire paper mentioning their mid-training process. The author notes that blending post-training data into mid-training—rather than just upsampling certain pretraining sources based on quality or grabbing domain-specific data—is becoming an industry standard practice.
More from Research
- Anthropic's Fable 5 Sets New High Score on ARC-AGI Benchmarks — mhmazur · 2026-08-06
- Specula: TLA+ Tool Automates Formal Specs, Finds Hundreds of Bugs — tianyin_xu · 2026-08-06
- Building Local AI NPC Systems with Emotion and Memory for Video Games — Patryk_Grzegorek · 2026-08-06
- UW professor Jerry Li wins 2026 Gödel Prize for solving robust statistics problem — lazowska · 2026-08-06
- Yi Ma Reaffirms Closed-Loop Feedback: End-to-End Will Return to Closed-Loop Learning — YiMaTweets · 2026-08-06
- 3DGS Meets Factor Graph SLAM: Unifying Pose Optimization and Rendering in GTSAM — fdellaert · 2026-08-06