IDC: global AI compute gap to hit $381B by 2030 as stacking GPUs stops working
量子位 · wechat · 2026-09-21
- IDC's report at AICC 2026 says global AI compute fulfillment will fall from 79% (2024) to 71% (2027), recovering only to 77% by 2030, with the absolute gap reaching $381B—about 10x 2024's level. Token consumption is forecast at a 4822.6% CAGR through 2030.
- Agent workloads turn compute demand from pulsed requests into continuous, multi-factor loads (task count x tokens per task x multi-agent amplification).
- Inspur CTO of AI Liu Jun frames the split as Capability vs Capacity compute, and launched two products:
- SD200Ultra: 128 chips, 8TB memory, 64TB storage; runs Kimi K3 (2.8T params) on a single node, supports up to 10T-param models; unified addressing and symmetric direct links cut All-to-All latency to 0.69 microseconds, and a "super operator" technique boosts K3 inference 3x.
- HC2000: liquid-cooled rack with >300kW per cabinet and 256 accelerator cards; heterogeneous chips split Prefill/Decode/FFN work, and the MCIS software stack delivers 10x token throughput at equal investment.
More from Infra
- Levelsio: A single $20 VPS once handled 300M visits a year — PratikKadam_ · 2026-09-21
- Qdrant experiments: 10→500 candidate depth lifts best-possible nDCG by 0.28 but real score by ≤0.01 — qdrant_engine · 2026-09-21
- gemini-cli PR Fixes Process Hang on Exit via stdin and MCP Child Process Cleanup — Pcmhacker-piro · 2026-09-21
- Daniel Lemire Tests Whether CPUs Can Take More Than One Branch Per Cycle — lemire · 2026-09-21
- Qwen3.8-27B in native 8-bit hits 37-55 tok/s on Apple Silicon, avoiding the 4-bit reasoning cliff — SnooPredictions515 · 2026-09-21
- Google's Orphaned VMs Patches Keep VMs Running While Host Kernel Goes Offline — jedisct1 · 2026-09-21