Zhipu Reveals Ox Alpha as GLM-5.3-Flash, Serving 100T Tokens/Day on Chinese GPUs
pstAsiatech · x · 2026-08-27
Zhipu AI has unmasked the mystery "Ox Alpha" model as GLM-5.3-Flash. The company revealed that the model runs entirely on domestic Chinese GPU hardware and currently processes 100 trillion tokens per day. This throughput figure is particularly striking given four years of chip export controls, highlighting the massive inference capabilities of China's local compute infrastructure.
More from Infra
- Concerns Arise Over HuggingFace's Hardware Neutrality After NVIDIA Acquisition — QuixiAI · 2026-08-27
- Nvidia posts blowout earnings; analyst calls it historic inflection point — PTrubey · 2026-08-27
- AI market success tied to Nvidia allocation, shifting partners to enemies — matt_slotnick · 2026-08-27
- Qwen Training Details: Muon Usage and TP Load Balancing — nrehiew_ · 2026-08-27
- Worried about HuggingFace? A Guide to Legally Back Up AI Models via Torrent — SeyAssociation38 · 2026-08-27
- Two vLLM recipes for Blackwell: NVFP4 KV cache buys 262K context and more streams — SeanHighness · 2026-08-27