HyperParallel FSDP2+Muon beats PyTorch FSDP2 throughput on Huawei Ascend cluster
PyTorch · x · 2026-09-09
PyTorch announces HyperParallel, a PyTorch-based distributed acceleration library optimized for Ascend SuperPoD. Under identical configurations, HyperParallel FSDP2 + Muon delivers substantially higher training throughput than PyTorch FSDP2 + Muon while maintaining consistent loss convergence.
The result was demoed live at the PyTorchCon China keynote on a Huawei Atlas 800T cluster, training Qwen3-30B-A3B with the complete distributed workflow, dynamically visualizing throughput and loss curves.
More from Infra
- Cerebras paper: layer dropout saves up to 25% training FLOPs and yields 1.55x faster decoding — burny_tech · 2026-09-09
- Viettel unifies GPU fleet into Token-as-a-Service platform with three open source layers — PyTorch · 2026-09-09
- Put per-turn action schemas in the last user message to preserve prompt caching — Low_Bad_6585 · 2026-09-09
- minnow: An Open-Source Fast Inference Server for LLaDA2.2 Diffusion LMs — coder543 · 2026-09-09
- PyTorch launches Accelerator Integration WG to fix hardware fragmentation — PyTorch · 2026-09-09
- HiSparse hybrid sparse attention lands in vLLM: 8x H200 concurrency jumps from 5 to 25 at 1M context — eliebakouch · 2026-09-09