Inspur Bets on Agent Infrastructure
量子位 · wechat · 2026-07-13
The core argument of this article is that the focus of AI infrastructure is shifting from "supporting LLM inference" to "enabling massive Agent scaling" and "continuously producing high-quality Tokens."
Background & Projections
- IDC projects China's enterprise-level AI Agent market to reach approximately 19 billion RMB in 2025, with a CAGR exceeding 110% from 2025 to 2028.
- Gartner estimates that by 2026, 40% of enterprise applications will integrate task-specific AI Agents.
- This means infrastructure must evolve beyond one-off input/output cycles to support task decomposition, tool calling, multi-turn collaboration, and persistent online availability.
Inspur Information's New Solutions
- Launched the industry's first CPU-native liquid-cooled full-rack server.
- A single rack supports up to 384 OCM-architecture CPUs, compatible with both x86 and ARM, enabling the concurrent operation of 40,000+ Agents.
- Adopts a "native liquid cooling" approach with co-designed compute and thermal management, extending liquid cooling to memory, NICs, optical modules, and SSDs.
- Improves space utilization and operational efficiency through 2U ultra-thin nodes, flat motherboard layouts, and a cable-free rack design, reportedly boosting O&M efficiency by 100%+.
Multi-Model Collaboration
- Launched a Multi-Modal Fusion API on the YuanBrain EPAI platform and released the YuanBrain SD200 Super-Node AI Server Enterprise Edition.
- The multi-modal fusion allows multiple candidate models to answer simultaneously, followed by a review fusion model that synthesizes consensus, disagreements, and omissions.
- Achieved 53.9% on the DRACO benchmark, outperforming any single model in the candidate pool.
- SD200's token generation time dropped from 8.9ms last year to 4.77ms, with first-token latency reduced by 35%.
Conclusion
The author concludes that infrastructure competition in the Agent era has shifted from "single-point model support" to "system-level synergy across CPUs, GPUs, and software platforms."
Related event: Inspur Unveils AI Infrastructure for the Agentic Era(2 posts)→
More from Infra
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22