Z.ai reveals Ox Alpha is GLM-5.3-Flash running on domestic chips

rohanpaul_ai · x · 2026-08-27

Z.ai revealed that Ox Alpha is actually GLM-5.3-Flash, powered by tens of thousands of domestic Chinese accelerators rather than NVIDIA GPUs. The model uses a MoE architecture with 320B total params but only 18B active during inference. It achieves a 10x cost reduction and significant benchmark improvements (DeepSWE 46.2->63.4) over GLM-5.2, utilizing linear attention to cut compute and KV cache size.

Related event: Mystery Model Ox Alpha Revealed as Zhipu's GLM-5.3-Flash Running on Chinese Chips(17 posts)→

Original post →

More from Infra

Infra channel →