Inside Zhipu's dense-feedback InfraAgent: an early blueprint for recursive self-improvement

量子位 · wechat · 2026-09-17

QbitAI relays Zhipu chief scientist Tang Jie's technical blog in full, detailing how a GLM-5.3-driven InfraAgent helped build and optimize a production inference system on 100k+ domestic accelerator cards. GLM-5.3-Flash ran anonymously as Ox-Alpha on OpenCode and OpenRouter, handling 62+ trillion tokens in six days.

The core technique is "dense feedback": organizing correctness tests, execution traces, and micro-benchmarks into the Agent's iteration loop, with feedback that is local, timely, and objectively verifiable—solving why end-to-end metrics alone can't explain regressions. Three case studies:

Zhipu stresses RSI hasn't been achieved—humans still set goals, boundaries, and review high-risk changes—but each Agent-completed engineering task may become training data for the next model generation.

Related event: Zhipu Discloses RSI Minimal Loop: GLM-5.3 Agent Built 100K-Card Domestic Inference Stack in Two Weeks(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →