Zhipu officially launches GLM-5.3-FlashX, its fastest model at 200 tokens/s
智谱 · wechat · 2026-09-18
Zhipu officially launched GLM-5.3-FlashX, its fastest model yet with inference speeds up to 200 tokens/s, now available via API (model key: GLM-5.3-FlashX).
The predecessor GLM-5.3-Flash — previously introduced to global developers under the alias "OxAlpha" — saw steadily growing adoption. FlashX builds on inference infrastructure powered by 100,000 domestic chips with additional infra-side optimization. Zhipu claims GLM-5.3-Flash remains the strongest intelligence in its size class at high cost-efficiency, and FlashX completes the package on intelligence, pricing, and speed.
More from Models
- openjev reranker based on Qwen3.5-4B trends on Hugging Face — AlexWortega · 2026-09-18
- User complains Codex hasn't reset for 6 days, pleads with Anthropic — cneuralnetwork · 2026-09-18
- Jev playground shows inference and roundtrip latency; EU users pay 120ms extra — DanielLockyer · 2026-09-18
- ChatGPT has stopped searching and started guessing: car repair data is all made up — WeissMISFIT · 2026-09-18
- Tesla publishes 12-month FSD Supervised safety report card — yunta_tsai · 2026-09-18
- ChatGPT co-inventor launches Jev model claiming 20-200x speed and 40-400x cost cuts — multiply_matrix · 2026-09-18