DeepSeek open-sources six Huawei Ascend projects, kernels hit 99.8% of chip peak
rohanpaul_ai · x · 2026-10-01
DeepSeek open-sourced six Huawei Ascend projects whose matrix kernels reach 99.8% of the chip's peak speed, further reducing reliance on Nvidia.
- The headline project is TileLang, the DSL DeepSeek writes its GPU kernels in; most operators used to train its V4 models now have Ascend versions, so the same source compiles for Huawei chips instead of Nvidia.
- DeepGEMM-Ascend reports 431 BF16 TFLOPS against a 432 ceiling.
- DeepEP-Ascend, which moves data between chips, shipped without a license file (the rest are MIT), so it can't be legally reused yet.
- Liang Wenfeng told investors Huawei could start delivering training chips this quarter.
Related event: DeepSeek Open-Sources Six Ascend Infrastructure Tools, Taking Aim at CUDA(17 posts)→
More from Infra
- Arduino argues sub-$900 embedded boards beat Mac minis for Physical AI agent economics — CatAstro_Piyush · 2026-10-01
- SemiAnalysis injects failures to rate GPU cluster renters — providers differ sharply on recovery — AccBalanced · 2026-10-01
- Redditor runs unattended DeepSeek loops for days: 237M tokens for just $3.48 — dogfoodarchitect · 2026-10-01
- Free course built from Cornell's GPU architecture workshop now shared publicly — idanbeck · 2026-10-01
- Qwen Flash Next MTP work resumes with official GGUF quants and llama.cpp PR — jacek2023 · 2026-10-01
- Undocumented Strata tip: set default sampling params via a sampling block in run config — KissMyShinyArse · 2026-10-01