Zhipu Reveals Ox Alpha as GLM-5.3-Flash, Serving 100T Tokens/Day on Chinese GPUs

pstAsiatech · x · 2026-08-27

Zhipu AI has unmasked the mystery "Ox Alpha" model as GLM-5.3-Flash. The company revealed that the model runs entirely on domestic Chinese GPU hardware and currently processes 100 trillion tokens per day. This throughput figure is particularly striking given four years of chip export controls, highlighting the massive inference capabilities of China's local compute infrastructure.

Related event: Mystery Model Ox Alpha Revealed as Zhipu's GLM-5.3-Flash Running on Domestic Chinese Chips(6 posts)→

Original post →

More from Infra

Infra channel →