Zhipu's InfraAgent boosted inference throughput 3.2x on 100k domestic chips in two weeks

APPSO · wechat · 2026-09-17

Zhipu chief scientist Tang Jie shared how GLM-5.3-driven InfraAgent helped deploy and optimize GLM-5.3-Flash on a cluster of over 100,000 domestic AI accelerators in just two weeks, raising end-to-end throughput 3.2x with per-token costs approaching mainstream NVIDIA GPU platforms.

The model itself performed much of the optimization work via a "dense feedback" mechanism: it located a TF32 precision-accumulation bug in KDA context-parallel paths (fix merged as FlashLinearAttention PR#1180), found that unreleased Python GIL in DeepEP caused over 30% KVTransfer overhead (reduced to under 1%), and lifted a redundant Decode Kernel by 1.71x by learning optimization patterns from SGLang and other open-source projects.

Zhipu frames this as an early form of recursive self-improvement: the model optimizes the system that serves the model, with engineers shifting toward designing feedback systems. The model previously ran anonymously as Ox-Alpha on OpenCode and OpenRouter, processing over 62 trillion tokens in six days.

Related event: Zhipu Discloses RSI Minimal Loop: GLM-5.3 Agent Built 100K-Card Domestic Inference Stack in Two Weeks(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →