Ox Alpha Revealed as Zhipu's GLM-5.3-Flash, Running on Chinese Chips All Along
eyishazyer · x · 2026-08-27
The mystery model Ox Alpha has been confirmed by Bloomberg as Zhipu's GLM-5.3-Flash. The poster spent six days piecing it together from tokenizer leaks and weird error codes before Bloomberg confirmed it outright, with weights dropped that evening.
The specs are aggressive: 320B total parameters with only 18B active, 1M token context, native text/image/video multimodality, MIT licensed, priced at $0.15 in / $0.50 out per million tokens. The arguably bigger headline: the whole stealth week ran on Chinese AI chips, not Nvidia. Benchmarks show real gains over GLM-5.2 and genuine competitiveness with Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash — though exact deltas are contested (a viral 10-task DeepSWE sample claimed 80%). The thread also previews a stacked release run: Qwen's Paloma, a possible Opus 5.1, Astra, and Grok 4.7.
More from Models
- 8x RTX 3090 Setup Serves Qwen Flash Next at 661 tok/s with 262k Context — QuixiAI · 2026-08-27
- GLM-5.3 Flash Slammed: Unstable, 'Schizophrenic' Behavior — kevinnbass · 2026-08-27
- Gemini Omni 1.1 Flash Spotted in Google Cloud Quotas, Launch May Be Near — koltregaskes · 2026-08-27
- Qwen3.8-Flash launches on OpenRouter with long video and multimodal support — Alibaba_Qwen · 2026-08-27
- Meme: Quantized Qwen 3.5 feels like an anxious tiny man — arthurcolle · 2026-08-27
- GLM-5.3 Flash tops OpenRouter share with 1% frontier cost on Chinese chips — jietang · 2026-08-27