Viral 'Ox Alpha' model revealed as GLM-5.3-Flash, served entirely on China-made chips
DeepLearningAI · x · 2026-09-08
DeepLearning.AI's The Batch reveals that the mystery model topping OpenRouter for over a week — 'Ox Alpha' — is Z.ai's GLM-5.3-Flash, with weights now downloadable. Key facts:
- Hybrid MoE transformer mixing linear and sparse attention: 320B total parameters, 18B active per token
- Native vision-language model (input up to 1,048,576 tokens, output 128,000 tokens at 44.6 tokens/sec) with tool calling, context caching, and adjustable reasoning levels that can't be disabled
- 57 on Artificial Analysis' Intelligence Index; best open-weights and third overall on GDPval-AA v2 at just $0.09 per task
- Available via GLM Coding Plan at $18–$168/month; the free high-volume preview ran entirely on China-made chips, with the company crediting memory optimization for overcoming hardware limits
More from Infra
- SemiAnalysis: TPU v7 Ironwood beats Blackwell Ultra by 50% perf per dollar on inference — Sentdex · 2026-09-08
- Intel extends High NA EUV lead as TSMC and Samsung confirm adoption by end of decade — BenBajarin · 2026-09-08
- User seeks a custom GGUF quant to fit GLM on a 192 GB RAM Mac between Q2 and Q4 — CentrifugalMalaise · 2026-09-08
- Community GPU guide compares GB per dollar and memory bandwidth across cards — jacek2023 · 2026-09-08
- Report claims 1 million high-NA EUV wafers per year, industry insider says it's possible — pstAsiatech · 2026-09-08
- Qwen3-0.6B (400MB) on a 2017 Galaxy Note 8 Drives Real Desktop Chrome via Structured Page Perception — Mean-Standard7390 · 2026-09-08