DeepSeek V4.1 Flash post-training details leak: synthetic data drives an RL revival

joecole · x · 2026-09-10

An unverified thread on DeepSeek V4.1 Flash's post-training highlights: vanilla SFT/RL/OPD algorithms with the focus on data synthesis; two synthesis routes (general agents and coding agents built from real repos, MCP/SWE-bench style); calibration across problems, envs and verifiers for correctness and difficulty, with the model itself trained to generate tasks; reasoning effort as a scalar with effort-wise grouping and credit assignment; elastic compute for sandboxing; and exponential length penalties tied to reasoning effort. Lots of infra detail on both pre- and post-training sides.

Original post →

More from Models

Models channel →