Pure CPU Runs 671B Model: 16 Supercomputer Nodes Match 80 GPUs in Throughput

新智元 · wechat · 2026-08-07

Qingcheng Jizhi and the domestic supercomputer "Lingshuang" have released a distributed inference solution for MoE models running purely on CPUs. Through deep software-hardware co-optimization, just 16 supercomputer nodes can smoothly run the DeepSeek-V3/R1-671B models.

At a batchsize=2048, the output throughput of this setup is comparable to an 80-mainstream-GPU cluster. On the hardware side, the "Lingshuang" supercomputer utilizes the self-developed LX2 CPU, which integrates on-chip high-bandwidth memory to provide TB-level bandwidth. This is paired with the "Lingqi" microsecond-level low-latency network, overcoming traditional CPU communication and memory bottlenecks.

On the software side, the engineering team rebuilt low-level operators specifically for the domestic CPU architecture and implemented a multi-level hybrid parallel strategy (e.g., large-scale EP256 expert parallelism). Combined with MTP (Multi-Token Prediction) technology, the single-request generation rate is significantly boosted. This deployment marks a shift in domestic computing power from merely competing on hardware specs to delivering cost-effective, commercially viable "Token factory" mass production.

Original post →

More from Infra

Infra channel →