Zhipu's GLM-5.3 Runs Entirely on Domestic Chips, Processing 62 Trillion Tokens
pstAsiatech · x · 2026-08-27
Zhipu AI launched its latest open-weight model, GLM-5.3-Flash (Ox Alpha), revealing it ran entirely on a cluster of 100,000 domestically produced chips during a stealth trial, processing 62 trillion tokens. It processed over 11 trillion tokens in its first three days on OpenRouter, becoming the platform's biggest launch and topping the coding charts. This marks a significant test of China's ability to handle large-scale inference on home-grown hardware.
More from Infra
- Cloudflare launches Computer: a virtual filesystem for AI agents — craigsdennis · 2026-08-27
- Pushing for LoRA sharing to reduce download waste — Borkato · 2026-08-27
- AI's Memory Crunch Hits Android Apps with New Limits — TechCrunch AI · 2026-08-27
- Local Deployment of GLM-5.3-Flash: 206 tok/s and 1M Context on DGX Station — funding__secured · 2026-08-27
- Hugging Face launches Jobs: run UV/Docker workloads on any hardware, pay per second — _akhaliq · 2026-08-27
- Merge, PostHog, and Redis host NYC technical talks on self-driving AI products — shensi · 2026-08-27