GLM-5.3-Flash handles 100T tokens daily, running entirely on Chinese chips
airesearch12 · x · 2026-08-26
SemiAnalysis reveals that Ox Alpha is actually Zhipu's GLM-5.3-Flash. Shockingly, the massive throughput of 100T tokens per day is served entirely on Chinese chips, highlighting the scalability of domestic silicon.
More from Infra
- Turso Adopts AgentID to Grant AI Agents Independent OIDC Identities — glcst · 2026-08-27
- Idea: Blockchain-based prompt credentialing for AI models — Dsphar · 2026-08-26
- Unsloth Releases Optimized Qwen2.5 Support for Faster Inference — thoquz · 2026-08-26
- Foresight CEO: Open Science Needs Independent Secure Compute Clusters — allisondman · 2026-08-26
- Qwen3.8-Flash Runs Locally: 125B Model on Just 75GB RAM — danielhanchen · 2026-08-26
- CoreWeave Powers Physical AI for Wayve, Decart, NEURA, and Nissan — wandb · 2026-08-26