Ascend SuperPOD optimization lifts DeepSeek-V4 post-training MFU to 34.22%
pmttyji · reddit · 2026-07-23
Full-parameter post-training on Ascend SuperPOD reaches 34.22% MFU
A paper on SLAI T-Rex describes an end-to-end optimization stack for full-parameter post-training of the trillion-parameter DeepSeek-V4 family on Ascend NPU SuperPOD infrastructure.
- The system tackles memory pressure, communication overhead, and kernel inefficiency across model parallelism, computation/communication orchestration, and low-level kernels.
- It reports 34.22% MFU, a 2.93× improvement over the open-source baseline recipe, while keeping training stable.
- The team also builds CPT/SFT pipelines for operations research tasks using DeepSeek-V4-Flash, combining collected domain resources with solver-verified synthetic documents.
- The resulting dataset has 10K SFT samples across four task categories and three problem representations.
- On the evaluated benchmarks, the specialized model reaches 71.81% zero-shot Pass@1, beating GPT-5.4-Mini by 3.98 points and the base DeepSeek-V4-Flash by 11.27 points.
The repo and paper present this as a full-stack path from infrastructure tuning to domain-specialized reasoning models.
Related event: Ascend SuperPOD Optimization Boosts DeepSeek-V4 Training MFU to 34.22%(2 posts)→
More from Infra
- AI video dubbing costs about $5–7 per finished minute once lip sync is included — Madmahi25 · 2026-07-23
- Nebius shows its first NVIDIA Vera Rubin NVL72 rack in Finland — demian_ai · 2026-07-23
- A Bittensor subnet launches inference at roughly half the usual price — markjeffrey · 2026-07-23
- Turbopuffer halves queue time after fixing deceptively hard autoscaling — DanielLockyer · 2026-07-23
- For a 1 GW data center, build 2 GW into the grid — anderssandberg · 2026-07-23
- Google is spending $200B+ on cloud and compute, Beff Jezos says — beffjezos · 2026-07-23