DeepSeek-V4 post-training on Ascend SuperPOD reaches 34.22% MFU
_akhaliq · x · 2026-07-23
SLAI T-Rex: DeepSeek-V4 post-training on Ascend SuperPOD
The paper describes an end-to-end post-training practice for trillion-parameter MoE models on Huawei Ascend SuperPOD.
- It targets DeepSeek-V4-family models and tackles the usual large-scale training pain points: memory pressure, communication overhead, and inefficient kernel execution.
- The authors build a hierarchical optimization stack across model parallelism, computation/communication orchestration, and low-level kernel execution.
- On the reported setup, the optimized system reaches 34.22% MFU, a 2.93× improvement over the open-source baseline while maintaining training stability.
- The same infrastructure is then used to build CPT and SFT pipelines for complex reasoning and OR tasks.
- They also introduce SLAI T-Rex with DeepSeek-V4-Flash, combining collected domain resources with solver-verified synthetic optimization documents.
- The dataset includes 10K high-quality SFT samples across four task categories and three problem representations.
- On the evaluated benchmark set, the model achieves the best average zero-shot Pass@1, reaching 71.81%, and beats GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 points, respectively.
Related event: Ascend SuperPOD Achieves 34.22% MFU for DeepSeek-V4 Training(3 posts)→
More from Infra
- AI agents are turning web search into a new infrastructure market — Annual_Judge_7272 · 2026-07-23
- Google’s raised $200B capex guide still failed to reassure AI investors — JOBhakdi · 2026-07-23
- vLLM says prime-rl runs trillion-scale agentic RL on 28 H200 nodes — vllm_project · 2026-07-23
- GLM-5.2 adds vision support and is now open source, with SGLang run instructions — baseten · 2026-07-23
- Atomic says its quantized-model runtime cuts KV cache use by up to 6.4x — testingcatalog · 2026-07-23
- DeepSeek says it has about 20,000 H-equivalent compute cards and will keep buying NVIDIA GPUs — ShakeelHashim · 2026-07-23