DeepSeek open-sources DeepGEMM Ascend port, hitting 99.8% of hardware limit on GEMM
zheanxu · x · 2026-09-30
DeepGEMM Ascend is now open source: DeepSeek's official port of DeepGEMM to Huawei Ascend NPUs, reportedly reaching 99.8% of the hardware limit on GEMM and 98% on MegaMoE.
Key points:
- Fully API-compatible with DeepGEMM; supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE
- Same APIs and dev workflow as DeepGEMM on other platforms after a simple install
- Lightweight abstraction over Ascend MAD primitives hides fractal layouts, alignment constraints, and address calculations
- Uses Ascend-specific optimizations like sparse data loading and coroutine-based pipelining to approach hardware limits
A landmark signal of DeepSeek embracing the domestic compute ecosystem.
More from Infra
- Raja Koduri: Western AI compute costs $50-60B per gigawatt, China targets under $10B — RajaXg · 2026-09-30
- AI intelligence-cost Pareto frontier shifted fast: GPT-5 mini at 17 ($0.05) to Claude Opus 5.5 at 58 ($5.98) — ArtificialAnlys · 2026-09-30
- Google's Project Suncatcher to fly TPUs in space for the first time on Oct 1 — allisondman · 2026-09-30
- Photon 2.6 ships FP8 + speculative decoding, runs Qwen3.5 27B at 400+ tok/s on B200 — Bedrovelsen · 2026-09-30
- SGLang turns Qwen3.8-27B into a decision model that beats Pokémon FireRed at sub-100ms — zhaoran_wang · 2026-09-30
- Agentic AI turns CPUs into the overlooked bottleneck as CPU:GPU ratios shift upward — AccBalanced · 2026-09-30