Ascend SuperPOD optimization lifts DeepSeek-V4 post-training MFU to 34.22%
pmttyji · reddit · 2026-07-23
Full-parameter post-training on Ascend SuperPOD reaches 34.22% MFU
A paper on SLAI T-Rex describes an end-to-end optimization stack for full-parameter post-training of the trillion-parameter DeepSeek-V4 family on Ascend NPU SuperPOD infrastructure.
- The system tackles memory pressure, communication overhead, and kernel inefficiency across model parallelism, computation/communication orchestration, and low-level kernels.
- It reports 34.22% MFU, a 2.93× improvement over the open-source baseline recipe, while keeping training stable.
- The team also builds CPT/SFT pipelines for operations research tasks using DeepSeek-V4-Flash, combining collected domain resources with solver-verified synthetic documents.
- The resulting dataset has 10K SFT samples across four task categories and three problem representations.
- On the evaluated benchmarks, the specialized model reaches 71.81% zero-shot Pass@1, beating GPT-5.4-Mini by 3.98 points and the base DeepSeek-V4-Flash by 11.27 points.
The repo and paper present this as a full-stack path from infrastructure tuning to domain-specialized reasoning models.
Related event: Ascend SuperPOD Achieves 34.22% MFU for DeepSeek-V4 Training(3 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11