iFlytek's Spark X2.5 trained on 10,000 domestic Ascend 910B GPUs with 97% uptime
机器之心 · wechat · 2026-09-11
A deep dive into how iFlytek trained Spark X2.5 (MoE, 293B-A30B) on a 10,000-GPU domestic Ascend 910B cluster, detailing the four core challenges of large-scale domestic compute:
- Compute efficiency: targeted Attention kernel optimization delivering 2x+ speedup over baseline
- Memory: sparse attention with cross-layer shared indexes plus dynamic load balancing for long sequences (+30% training efficiency)
- Communication: compression, hierarchical communication and compute parallelization (+10% efficiency)
- Stability: pre-run validation, automatic hang detection, failed-task resubmission and process-level recovery, achieving >97% effective training time
The model uses multi-teacher online policy distillation (MOPD) to boost coding and agent capabilities. In one case, Spark X2.5 autonomously explored 1,600 tool calls to optimize the DSA sparse-attention kernel, achieving 3.5x speedup over torchnpu — showing models starting to optimize the NPU software stack itself, forming a feedback loop between domestic compute and model training. Adaptation work is also underway on Cambricon MLU590 and Hygon BW1000.
More from Infra
- Web Demo Approximates V4.1 Flash-Style Fast KV Prefill on Qwen3 — T_rex2700 · 2026-09-11
- OpenAI CFO: Compute I bought a year ago could sell for 3-5x today — and we're still short — rwang07 · 2026-09-11
- REVA Mines LLM Attention into Reusable Evidence Views, Cutting RAG Compression Overhead up to 15.6x — _reachsumit · 2026-09-11
- Edge0-35B-A3B preview MoE model for edge inference trends on Hugging Face — Edge0 · 2026-09-11
- Reflect Orbital wants to sell sunlight via volleyball-court mirrors on satellites — kyliebytes · 2026-09-11
- Data centers are for startups, not frontier labs: more compute is the anti-monopoly move — arthurcolle · 2026-09-11