Z.ai runs GLM-5.3-Flash entirely on Chinese AI chips, cutting costs by over 40%

yogthos · reddit · 2026-08-28

Z.ai has successfully migrated the inference of its GLM-5.3-Flash model entirely to domestic Chinese AI chips, eliminating reliance on NVIDIA hardware. Benchmarks indicate that the domestic chip solution reduces computing costs by more than 40% compared to using NVIDIA GPUs. This achievement validates the feasibility and economic viability of China's domestic compute infrastructure for large-scale LLM inference.

Original post →

More from Infra

Infra channel →