DeepSeek's near-hardware-limit kernels hint at a new model, possibly Ascend-trained
teortaxesTex · x · 2026-09-30
X user teortaxesTex argues DeepSeek would not have achieved 99.8% of the hardware limit on GEMM and 98% on MegaMoE and open-sourced those kernels without training something usable end to end — the kernel release likely supports post-training for an upcoming model. He speculates the rumored small single-GPU model may be Ascend-born. Unverified speculation.
Related event: DeepSeek's New Model May Be Trained on Huawei Ascend Chips(3 posts)→
More from Infra
- Quantized softmax attention pretraining: only +0.004 nats loss gap at K=16 with the right calibration — illinois · 2026-09-30
- xLLM training infra open-sourced with xattn attention backend and xBridges toolkit — HongyiWang10 · 2026-09-30
- Auto-research loop on 120 B300s finds 40% Kimi K3 inference gain for $9,176 — bookwormengr · 2026-09-30
- Cerebras to bring 'world's fastest inference' to General Compute — beffjezos · 2026-09-30
- 8% of Asia-to-US air freight is now data center parts — 30 full freighters a day — yacineMTB · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30