Inspur Unveils Agent-Era AI Compute Infrastructure

机器之心 · wechat · 2026-07-13

At the Open Compute Project conference, Inspur unveiled a suite of AI infrastructure products tailored for the agentic era, featuring **CPU-native liquid-cooled full-rack servers** and **YuanNao SD200 ultra-node AI servers**. The core thesis is that compute architecture is shifting as AI transitions from single-turn Q&A to collaborative multi-agent systems. CPUs are evolving beyond simple scheduling to handle workloads like agent sandboxes, workflow orchestration, and tool invocations. Consequently, data centers require purpose-built **CPU compute foundations** alongside GPUs, which continue to manage large model inference. On the hardware front, the CPU-native liquid-cooled racks utilize an open OCM architecture, emphasizing high density, full-path liquid cooling, and elevated rack power limits to support massive agent deployments. On the software and inference side, the upgraded YuanNao SD200 ultra-node reportedly reduces the single Token generation time to **4.77ms** for the **Kimi K2.6 trillion-parameter model**, while also lowering the time-to-first-Token. It delivers high-performance optimizations for open-source models like Kimi, DeepSeek, GLM, and MiniMax. The article concludes by summarizing the architecture: **GPU ultra-nodes handle the "thinking," while CPU-native liquid-cooled racks manage the "acting"**, jointly driving the enterprise adoption of agents and multimodal applications.

Related event: Inspur Unveils AI Infrastructure for the Agentic Era(2 posts)→

Original post →

More from Infra

Infra channel →