Zhipu's InfraAgent boosted inference throughput 3.2x on 100k domestic chips in two weeks
APPSO · wechat · 2026-09-17
Zhipu chief scientist Tang Jie shared how GLM-5.3-driven InfraAgent helped deploy and optimize GLM-5.3-Flash on a cluster of over 100,000 domestic AI accelerators in just two weeks, raising end-to-end throughput 3.2x with per-token costs approaching mainstream NVIDIA GPU platforms.
The model itself performed much of the optimization work via a "dense feedback" mechanism: it located a TF32 precision-accumulation bug in KDA context-parallel paths (fix merged as FlashLinearAttention PR#1180), found that unreleased Python GIL in DeepEP caused over 30% KVTransfer overhead (reduced to under 1%), and lifted a redundant Decode Kernel by 1.71x by learning optimization patterns from SGLang and other open-source projects.
Zhipu frames this as an early form of recursive self-improvement: the model optimizes the system that serves the model, with engineers shifting toward designing feedback systems. The model previously ran anonymously as Ox-Alpha on OpenCode and OpenRouter, processing over 62 trillion tokens in six days.
More from AGI Musings
- Polymarket puts 36% odds on frontier AI labs agreeing to pace AI by 2026 — Polymarket · 2026-09-17
- A Private WoW Server Would Be the Perfect Sandbox for Testing AI General Intelligence — djcows · 2026-09-17
- A 2008 Story Description Now Reads Like LLM Output: 'Semantic Apocalypse' — erikphoel · 2026-09-17
- OpenAI cracks a math logjam as 25 Fields medalists sign cautionary letter — nordicinst · 2026-09-17
- 'By 2030' AI safety assurances don't reassure, commentator notes — FlorianGallwitz · 2026-09-17
- Nate Silver voices skepticism on RSI, spars with AI class-action plaintiff — dhadfieldmenell · 2026-09-17