Zhipu Discloses Minimal RSI Loop: GLM-5.3 Agent Built Inference Stack on 100K Domestic GPUs in Two Weeks
Zhipu founder and Chief Scientist Tang Jie posted a long thread on X and shared a technical blog, disclosing for the first time the company's minimal closed-loop result in recursive self-improvement (RSI): InfraAgent, powered by GLM-5.3, helped build the production-grade inference service for GLM-5.3-Flash on a cluster of over 100,000 domestic AI accelerators, going from the model's first successful run on domestic accelerators to handling all production traffic in less than two weeks, with end-to-end throughput reaching 3.2x the baseline. This is a public demonstration of "a model optimizing its own inference infrastructure" and a landmark case for bringing a large-scale domestic chip cluster online.
Confirmed
- All production inference traffic for GLM-5.3-Flash now runs on over 100,000 domestic AI accelerators, with less than two weeks from first successful run to production launch
- End-to-end throughput improved 3.2x over baseline; optimization metrics such as per-token cost are disclosed in Tang Jie's blog (as relayed by @APPSO)
- Much of the optimization work was done by the GLM-5.3-powered Infra Agent, meaning the model participated in optimizing the infrastructure that carries inference for its own model family, forming a minimal RSI loop
- According to @量子位's detailed analysis, the system's technical approach centers on dense feedback and is seen as a prototype of RSI; GLM-5.3-Flash was previously available under the anonymous name Ox-Alpha on platforms like OpenRouter
Why it matters
- This is the first publicly disclosed case in China of "using one's own model to optimize one's own inference infrastructure at scale," validating RSI's feasibility in a real production environment rather than just on paper
- A 100,000-card domestic accelerator cluster handling all production traffic is a real-world test for both the domestic compute supply chain and inference system optimization capabilities
- First-hand accounts from Zhipu engineer jietang show the two-week sprint from first run to full deployment, with Infra Agent handling most optimizations — a human-machine division of labor that offers a reference for the industry
2026-09-17 ~ 2026-09-17 · 8 related posts
Primary sources
- Zhipu Details How GLM Built Its Own Inference Infrastructure Toward Recursive Self-Improvement — likeastar20 ·
- GLM agent built its own inference infra in two weeks, tripling end-to-end throughput — jietang ·
- Zhipu's InfraAgent boosted inference throughput 3.2x on 100k domestic chips in two weeks — APPSO ·
- [source] GLM agent built its own inference infra in two weeks, tripling end-to-end throughput — jietang · 2026-09-17
- Zhipu says a GLM-5.3-powered agent built production inference infra in under two weeks — 智东西 · 2026-09-17
- Inside Zhipu's dense-feedback InfraAgent: an early blueprint for recursive self-improvement — 量子位 · 2026-09-17
- Zhipu says a GLM-5.3 agent optimized its own inference stack, 3.2x throughput on 100k domestic chips — SinclairWang1 · 2026-09-17
- [source] Zhipu's InfraAgent boosted inference throughput 3.2x on 100k domestic chips in two weeks — APPSO · 2026-09-17
- GLM-5.3-powered Infra Agent boosts own inference stack 3x on 100k Chinese accelerators — Dr_Singularity · 2026-09-17
- Zhipu Claims GLM-Powered Infra Agent Delivered 3x Speedup on 100k+ Domestic Accelerators — Dr_Singularity · 2026-09-17
- [source] Zhipu Details How GLM Built Its Own Inference Infrastructure Toward Recursive Self-Improvement — likeastar20 · 2026-09-17