DeepSeek-V4 post-training on Ascend SuperPOD reaches 34.22% MFU
Dongfang Li · hf · 2026-07-23
What it reports
This paper describes an end-to-end post-training optimization stack for the DeepSeek-V4 family on an Ascend NPU SuperPOD.
- It tackles trillion-parameter MoE post-training bottlenecks such as memory pressure, communication overhead, and inefficient kernels.
- The authors build a hierarchical optimization framework across parallelism, compute/communication orchestration, and kernel execution.
- The system reaches 34.22% MFU, a 2.93× improvement over the open-source baseline, while keeping training stable.
- On top of that infrastructure, they create a CPT/SFT workflow for complex Operations Research tasks using solver-verified synthetic data.
- The resulting specialized model gets 71.81% zero-shot Pass@1, beating GPT-5.4-Mini by 3.98 points and the base DeepSeek-V4-Flash by 11.27 points.
More from Infra
- Google capex debate centers on data centers, not just cloud unit returns — aronchick · 2026-07-23
- OpenAI’s planned Australian data center drops recycled-water cooling as Sydney grid nears capacity — Polymarket · 2026-07-23
- Most tasks can run on cheap models, so smart routing beats using Opus everywhere — bindureddy · 2026-07-23
- Four RTX 3080s hit 69 tok/s on Qwen3.6-27B for about $2,000 — starkruzr · 2026-07-23
- Free market-data API ships with llms.txt, OpenAPI and an Agent Skill — CaseLivid4116 · 2026-07-23
- Oriental Computing’s DF1000 is pitched as a 14nm chip with Hopper-like utility — teortaxesTex · 2026-07-23