Zhipu's GLM 3.5 Flash Served 42T Tokens in 6 Days Free on Chinese Chips
bindureddy · x · 2026-08-27
A breakdown of how Ox-Alpha (GLM 3.5 Flash) processed 42 trillion tokens for free over 6 days, entirely on Chinese silicon.
Key points:
- Custom SGLang inference engine with disaggregated Encode-Prefill-Decode architecture
- A GLM-5.3 infrastructure agent wrote its own GPU kernels and debugged bottlenecks — the model optimized its own serving stack
- 3× end-to-end performance gains, Nvidia-level cost parity on domestic silicon
- Backed by Zai's 1GW data center: 10,000+ chip clusters, zero Nvidia silicon
The author argues export controls backfired, forcing a fully sovereign Chinese AI stack that works.
More from Infra
- Optical connectivity standards: CPO has rules, but NPO is chaos — jwt0625 · 2026-08-27
- ABF Substrates & PCB Identified as Key Constraints for 2027 AI Hardware — zephyr_z9 · 2026-08-27
- First startup enriches uranium for nuclear-powered data centers — Polymarket · 2026-08-27
- G2 spent $1.27M on 970B tokens, shifting focus from adoption to efficiency — prasanna_says · 2026-08-27
- TokenVisor supports Nvidia, AMD, and Intel GPUs in a single cluster — AccBalanced · 2026-08-27
- Zai's domestic inference cluster hits 100k+ chips; GLM-5.3 runs on custom interconnect — zephyr_z9 · 2026-08-27