Mystery Model Ox Alpha Revealed as Zhipu's GLM-5.3-Flash Running on Domestic Chinese Chips
The mystery model Ox Alpha has been revealed to be Zhipu's GLM-5.3-Flash. Bindu Reddy first spotted the connection on August 26, and Zhipu subsequently confirmed it, stating that the model runs entirely on Chinese-made GPUs and claims to handle 100 trillion tokens of daily calls. The unusual scale of free access and the possibility of a domestic compute foundation have drawn intense community attention.
Confirmed
- Ox Alpha is Zhipu's GLM-5.3-Flash, discovered by Bindu Reddy and confirmed by Zhipu.
- GLM-5.3-Flash features a 1 million context window.
- Zhipu claims the model handles up to 100 trillion (100T) tokens per day and offers free access at massive scale.
- SemiAnalysis reported that the 100T tokens-per-day workload runs entirely on Chinese chips, demonstrating large-scale orchestration on a domestic compute base.
Not yet confirmed
- The compute backer is disputed: Bindu Reddy speculated that Nvidia may be providing the compute, enabling free access of up to 100T tokens for training/inference, calling it a "genius marketing strategy for open-source AI"; however, @AccBalulated's analysis suggests Huawei, not Nvidia, is the backer. Both remain speculation with no official confirmation.
- Community discussion of the exact scale of free access (estimated at 100 trillion tokens) and per-user quota limits also remains speculative.
Why it matters
- If the 100T tokens-per-day workload is indeed supported by domestic Chinese chips, it would be a public demonstration of China's AI compute capability in ultra-large-scale inference orchestration—especially significant after years of chip restrictions (@pstAsiatech referenced the "after 4 years of chip…" context).
- The massive free access itself is a major competitive signal: whether backed by Nvidia or Huawei, offering a top-tier model to developers at extremely low cost could reshape the open-source AI ecosystem.
2026-08-26 ~ 2026-08-27 · 6 related posts
Primary sources
- [source] Nvidia may have funded 100T tokens for free GLM 5.3 Flash release — bindureddy · 2026-08-26
- Speculation: Nvidia Powered GLM's 100T Token Free Giveaway — bindureddy · 2026-08-26
- GLM-5.3-Flash handles 100T tokens daily, running entirely on Chinese chips — airesearch12 · 2026-08-26
- Rumor: Huawei, Not Nvidia, Backed GLM 5.3 Flash's Massive Token Run — AccBalanced · 2026-08-27
- [source] Zhipu Reveals Ox Alpha as GLM-5.3-Flash, Serving 100T Tokens/Day on Chinese GPUs — pstAsiatech · 2026-08-27
- [source] Zhipu's GLM 3.5 Flash Served 42T Tokens in 6 Days Free on Chinese Chips — bindureddy · 2026-08-27