DeepSeek V4.1 Flash post-training details leak: synthetic data drives an RL revival
joecole · x · 2026-09-10
An unverified thread on DeepSeek V4.1 Flash's post-training highlights: vanilla SFT/RL/OPD algorithms with the focus on data synthesis; two synthesis routes (general agents and coding agents built from real repos, MCP/SWE-bench style); calibration across problems, envs and verifiers for correctness and difficulty, with the model itself trained to generate tasks; reasoning effort as a scalar with effort-wise grouping and credit assignment; elastic compute for sandboxing; and exponential length penalties tied to reasoning effort. Lots of infra detail on both pre- and post-training sides.
More from Models
- OpenAI developers site adds showcase page with GPT-6 Astra outputs — Dimillian · 2026-09-10
- Arena.ai: Claude Fable 5.1 Writes More Matter-of-Fact but More Verbose — The Decoder · 2026-09-10
- Qwen-Image-Edit-2511 vs SenseNova U1.5 Lite: hands-on multi-reference fusion comparison — daniel933912 · 2026-09-10
- Cognition launches SWE-2: frontier-level coding performance at up to 70% lower cost — silasalberti · 2026-09-10
- NeoHorse-1-4B, a Qwen3.5-based agentic model, trends on Hugging Face — TokenRhythm · 2026-09-10
- Chinese model 3D face-off: DeepSeek V4.1 Flash crushes Kimi K3 and GLM-5.3 — teortaxesTex · 2026-09-10