Zhipu officially launches GLM-5.3-FlashX, its fastest model at 200 tokens/s

智谱 · wechat · 2026-09-18

Zhipu officially launched GLM-5.3-FlashX, its fastest model yet with inference speeds up to 200 tokens/s, now available via API (model key: GLM-5.3-FlashX).

The predecessor GLM-5.3-Flash — previously introduced to global developers under the alias "OxAlpha" — saw steadily growing adoption. FlashX builds on inference infrastructure powered by 100,000 domestic chips with additional infra-side optimization. Zhipu claims GLM-5.3-Flash remains the strongest intelligence in its size class at high cost-efficiency, and FlashX completes the package on intelligence, pricing, and speed.

Original post →

More from Models

Models channel →