iFlytek's Spark X2.5 trained on 10,000 domestic Ascend 910B GPUs with 97% uptime

机器之心 · wechat · 2026-09-11

A deep dive into how iFlytek trained Spark X2.5 (MoE, 293B-A30B) on a 10,000-GPU domestic Ascend 910B cluster, detailing the four core challenges of large-scale domestic compute:

The model uses multi-teacher online policy distillation (MOPD) to boost coding and agent capabilities. In one case, Spark X2.5 autonomously explored 1,600 tool calls to optimize the DSA sparse-attention kernel, achieving 3.5x speedup over torchnpu — showing models starting to optimize the NPU software stack itself, forming a feedback loop between domestic compute and model training. Adaptation work is also underway on Cambricon MLU590 and Hygon BW1000.

Original post →

More from Infra

Infra channel →