Zhipu Open-Sources GLM-5.3-Flash: 320B MoE with 1M Context at a Fraction of the Cost
On August 26, Zhipu AI officially released and open-sourced GLM-5.3-Flash — the mystery model previously tested quietly on OpenRouter under the name Ox Alpha. The model's weights are available on Hugging Face under the MIT license, positioned as a combination of frontier performance and extremely low cost, and it sparked massive community discussion on launch day.
Confirmed
- Specs and architecture: 320B total parameters with 18B activated in a MoE architecture; the first open-weights release of the glm5next architecture; uses a hybrid of sparse and linear attention, is natively multimodal, supports a 1 million token context window; FP8 weights take up roughly 328GB.
- Performance: officially, it comprehensively surpasses GLM-5.2 on the Code Bench real-world coding benchmark and matches Claude Opus 4.8; third-party Artificial Analysis gives it an intelligence index of 57, placing it on the Pareto frontier alongside the full GLM-5.3; according to op7418, during its preview period on OpenRouter it already consumed 23.2T tokens, ending DeepSeek's streak at the top.
- Pricing: the official API is priced at $0.15 per million input tokens, $0.50 output, and $0.03 cached input. Multiple posts say that's roughly one-tenth to one-fifteenth of GLM-5.2 or the full GLM-5.3, and about 40x cheaper than Opus 4.8; a review relayed by Hesamation put single-task cost at around $0.045.
- Ecosystem: already supported by vLLM; the Hugging Face page is based on the Transformers framework.
- Compute: according to posts by orange and others, the company claims daily supply of over 100T tokens is backed by domestic (Ascend) chips; zephyrz9 says the model is deployed on Ascend hardware, with FP8 weights at about 328GB, and NVFP4 is not yet used.
Not Yet Confirmed
- Claims that "the global daily supply of over 100T tokens is entirely backed by domestic chips," along with some of the cost-performance multiples, come from official marketing or community relay, with no independent verification yet.
Why It Matters
- With less than half of GLM-5.2's parameter count yet comprehensively surpassing the previous generation across benchmarks, combined with aggressive pricing, it poses a direct challenge to the value proposition of closed frontier models; the hybrid sparse/linear attention route significantly cuts long-context serving costs and offers the open-source community a new architectural blueprint.
2026-08-25 ~ 2026-08-27 · 29 related posts
- Episode 1: Mystery model Ox Alpha stuns community; fingerprints point to Zhipu GLM-5.3(2026-08-21, 58 posts)
- Episode 2: OpenRouter's Mystery 'Ox Alpha' Model Identified as GLM(2026-08-21, 3 posts)
- Episode 3: Cola Launches Ox Model Claiming Terra-Level Capability at Flash Speed(2026-08-21, 2 posts)
- Episode 4: Mysterious Model Ox Alpha Gives Away 100T Tokens Daily, Sparking Power Questions(2026-08-21, 2 posts)
- Episode 5: Ox-Alpha Underperforms in Tests, Called Marketing Stunt(2026-08-22, 3 posts)
- Episode 6: Ox Alpha scores ~63% on full DeepSWE; earlier 80% was subset artifact(2026-08-22, 6 posts)
- Episode 7: Forensics Point to Zhipu's Unreleased GLM Behind Mystery Model Ox Alpha(2026-08-22, 31 posts)
- Episode 8: Ox Alpha Generates 64K-Token 3D Scene in One Prompt(2026-08-22, 2 posts)
- Episode 9: Mystery model Ox Alpha draws mixed reviews from early testers(2026-08-22, 6 posts)
- Episode 10: Mystery Model Ox Alpha Launches Free, Identity Unknown(2026-08-23, 3 posts)
- Episode 11: Mysterious Model Ox Alpha Debuts on OpenRouter, Sparking Speculation(2026-08-24, 5 posts)
- Episode 12: Leak: Ox Alpha Is GLM-5.3 Flash, 320B MoE at One-Tenth the Price(2026-08-25, 4 posts)
- Episode 13: Mystery chart-topper Ox Alpha confirmed as Zhipu GLM-5.3-Flash, weights open-sourced tonight(2026-08-25, 16 posts)
- Episode 14: Zhipu Open-Sources GLM-5.3-Flash: 320B MoE with 1M Context at a Fraction of the Cost(2026-08-25, 29 posts)
- Episode 15: Stealth Model Ox Alpha Proves Mediocre at Coding, Trails Grok 4.6(2026-08-25, 3 posts)
- Episode 16: Mystery Model Ox Alpha Burns 42T Tokens in Six Days(2026-08-26, 3 posts)
Primary sources
- [source] Zai Releases GLM-5.3-Flash Model on Hugging Face — zai-org · 2026-08-25
- Zhipu open-sources GLM-5.3-Flash: matches Claude Opus 4.8 at 1/40 the price — 智谱 · 2026-08-26
- [source] Zai Announces GLM-5.3-Flash Pricing: $0.15 per 1M Input Tokens — Zai_org · 2026-08-26
- [source] GLM-5.3-Flash Matches Claude Opus 4.8 on Code Bench — Zai_org · 2026-08-26
- GLM-5.3-Flash: Architecture Enhancements Boost Efficiency — Zai_org · 2026-08-26
- Zai Launches 320B Parameter GLM-5.3-Flash Model — TheZachMueller · 2026-08-26
- Zhipu releases GLM-5.3-Flash: 320B params, MIT open source — TheZachMueller · 2026-08-26
- Zhipu quietly releases glm-5.3-flash model — koltregaskes · 2026-08-26
- Zhipu's GLM-5.3-Flash Model Released on Hugging Face — coder543 · 2026-08-26
- Zhipu releases GLM-5.3-Flash: smaller size, Opus 4.8 level performance — airesearch12 · 2026-08-26
- GLM-5.3-Flash Uses Hybrid Attention to Cut Long-Context Costs — multimodalart · 2026-08-26
- GLM-5.3-Flash adopts sparse + linear attention hybrid architecture — multimodalart · 2026-08-26
- Zhipu Releases GLM-5.3-Flash: Frontier Intelligence, Flash Cost — dydynam · 2026-08-26
- Zhipu Launches GLM 5.3 Flash Vision: 320B Params at 1/40th of Opus 4.8's Price — oran_ge · 2026-08-26
- GLM-5.3 Flash: 320B total params, 18B active, deployed on Ascend — zephyr_z9 · 2026-08-26
- Zhipu Releases GLM 5.3: 100T Daily Compute on Domestic Chips, Costs Slashed by 90% — oran_ge · 2026-08-26
- GLM-5.3-Flash released with benchmark comparisons — elemental-mind · 2026-08-26
- Zhipu Releases GLM 5.3 Flash: 320B Params, 90% Price Cut — oran_ge · 2026-08-26
- Zai Releases GLM-5.3-Flash: 320B MoE with Hybrid Attention and 1M Context — TheZachMueller · 2026-08-26
- Zhipu Releases GLM-5.3-Flash: 320B MoE Open Source Under MIT — realsohamparekh · 2026-08-26
- GLM-5.3-Flash scores 57 on Artificial Analysis, hits Pareto frontier — ollama · 2026-08-26
- GLM-5.3-Flash matches Opus 4.8, 45x cheaper — Hesamation · 2026-08-26
- Zhipu GLM-5.3-Flash revealed: 320B params, MIT open source — op7418 · 2026-08-26
- Z.ai Releases GLM-5.3-Flash: Hybrid Sparse+Linear Attention Architecture — No_Afternoon_4260 · 2026-08-26
- New Qwen and GLM Open Models Benchmark Against Claude Opus — kimmonismus · 2026-08-26
- GLM-5.3-Flash Released: 320B MIT-Licensed Model Outperforms Predecessor — matei_zaharia · 2026-08-26
- Zhipu releases GLM-5.3-Flash: Frontier intelligence at flash cost — 1littlecoder · 2026-08-27
- Zhipu Open Sources GLM-5.3 Flash: Efficient 320B Model Ties Claude Opus — 量子位 · 2026-08-27
- GLM-5.3-Flash released: 320B MoE architecture, MIT licensed — TheZachMueller · 2026-08-27