DeepSeek open-sources full Ascend AI training stack, mirroring its NVIDIA components
aigclink · x · 2026-09-30
DeepSeek has open-sourced the entire in-house infrastructure stack it built for training its V4-series models, ported from NVIDIA GPUs to Huawei Ascend chips.
The stack powers most of the operators used in V4 training and includes the TileLang kernel DSL, the DeepGEMM matrix computation library, the DeepEP distributed communication library, plus TileKernels, FlashMLA, and DeepSelect — each component has a one-to-one NVIDIA counterpart. It marks the first time a full, open training stack of this scope is available for Ascend hardware.
Related event: DeepSeek Open-Sources Full Ascend AI Infrastructure Stack(5 posts)→
More from Infra
- Quantized softmax attention pretraining: only +0.004 nats loss gap at K=16 with the right calibration — illinois · 2026-09-30
- xLLM training infra open-sourced with xattn attention backend and xBridges toolkit — HongyiWang10 · 2026-09-30
- Auto-research loop on 120 B300s finds 40% Kimi K3 inference gain for $9,176 — bookwormengr · 2026-09-30
- Cerebras to bring 'world's fastest inference' to General Compute — beffjezos · 2026-09-30
- 8% of Asia-to-US air freight is now data center parts — 30 full freighters a day — yacineMTB · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30