Z.ai reveals Ox Alpha is GLM-5.3-Flash running on domestic chips
rohanpaul_ai · x · 2026-08-27
Z.ai revealed that Ox Alpha is actually GLM-5.3-Flash, powered by tens of thousands of domestic Chinese accelerators rather than NVIDIA GPUs. The model uses a MoE architecture with 320B total params but only 18B active during inference. It achieves a 10x cost reduction and significant benchmark improvements (DeepSWE 46.2->63.4) over GLM-5.2, utilizing linear attention to cut compute and KV cache size.
More from Infra
- Grokpute launches distributed GPU training network — jw2yang4ai · 2026-08-27
- Weaviate ships query profiling: one flag pinpoints slow-query bottlenecks inline — victorialslocum · 2026-08-27
- Storage architecture for AI sandboxes: Local root + S3 — aniketmaurya · 2026-08-27
- Weaviate adds 'effort' parameter to tune search quality vs. compute cost — CShorten30 · 2026-08-27
- Managing Token Spend in Multi-Agent Setups — souvlakee · 2026-08-27
- Organizations Spend Over $117k Monthly on Agentic AI Inference — perilli · 2026-08-27