Pure CPU Runs 671B Model: 16 Supercomputer Nodes Match 80 GPUs in Throughput
新智元 · wechat · 2026-08-07
Qingcheng Jizhi and the domestic supercomputer "Lingshuang" have released a distributed inference solution for MoE models running purely on CPUs. Through deep software-hardware co-optimization, just 16 supercomputer nodes can smoothly run the DeepSeek-V3/R1-671B models.
At a batchsize=2048, the output throughput of this setup is comparable to an 80-mainstream-GPU cluster. On the hardware side, the "Lingshuang" supercomputer utilizes the self-developed LX2 CPU, which integrates on-chip high-bandwidth memory to provide TB-level bandwidth. This is paired with the "Lingqi" microsecond-level low-latency network, overcoming traditional CPU communication and memory bottlenecks.
On the software side, the engineering team rebuilt low-level operators specifically for the domestic CPU architecture and implemented a multi-level hybrid parallel strategy (e.g., large-scale EP256 expert parallelism). Combined with MTP (Multi-Token Prediction) technology, the single-request generation rate is significantly boosted. This deployment marks a shift in domestic computing power from merely competing on hardware specs to delivering cost-effective, commercially viable "Token factory" mass production.
More from Infra
- Open-Weight Small Models Enable Fully Local AI Tasks from ASR to Agents — vanstriendaniel · 2026-08-07
- SpaceX Generates Up to $50M/MW in Compute, Crushing Neoclouds — shaunmmaguire · 2026-08-07
- DeepInfra Processes Over 500B Tokens Daily Amid Surging Inference Demand — patricious · 2026-08-07
- AMD ROCm Open-Sources Spur: A Rust-Based, AI-Native GPU Scheduler — AnushElangovan · 2026-08-07
- Bun 1.4 Canary Slashes Idle CPU Usage Significantly — DanielLockyer · 2026-08-07
- Handling 9B Daily Requests: Cloudflare Migrates cdnjs to Its Developer Platform — JeremyCMorgan · 2026-08-07