Zhipu open-sources GLM-5.3-Flash, a 320B MoE frontier model
Zhipu officially released and open-sourced GLM-5.3-Flash (codename Ox Alpha) on August 26. It is the first natively multimodal model in the GLM-5 series and the first open-weights release of the glm5next architecture. Published on Hugging Face under the MIT license, it positions itself as "frontier intelligence + Flash speed at Flash cost," with the company claiming the price is just one-tenth of GLM-5.2's.
Confirmed
- Architecture & scale: A sparse MoE architecture with 320B total parameters and 18B activated parameters, combining Kimi Linear hybrid attention with DeepSeek-related techniques (sparse attention + linear attention); natively multimodal, supports a 1 million token context window, and already has vLLM support
- Pricing: Official standard API pricing is $0.15 per 1M input tokens, $0.50 per 1M output tokens, and $0.03 for cached input
- Performance: Officially, it significantly outperforms GLM-5.2 across all compute tiers on the Code Bench real-world coding benchmark and matches Claude Opus 4.8; it scores 57 on the Artificial Analysis intelligence index, only 3 points below the full GLM-5.3, sitting on the Pareto frontier alongside GLM-5.3
- Hardware deployment: According to @zephyrz9 and @orange, the FP8 weights are roughly 328GB and the model is deployed on domestic chips such as Ascend; the company says its daily supply of over 100T of compute is fully backed by domestic chips
Not Yet Confirmed
- Third-party replications of the benchmark data are still pending; the claimed performance comparisons (such as matching Opus 4.8) currently come mainly from the official blog and preliminary Artificial Analysis evaluations
Why It Matters
- With less than half the parameters of GLM-5.2 yet surpassing it across benchmarks, paired with ultra-low pricing, observers like @bindureddy see potential for it to become the most widely used open-source model; @Hesamation estimates its per-task cost at about $0.045 — 15x cheaper than GLM-5.3 and roughly 45x cheaper than Opus 4.8
- All inference compute is backed by domestic chips; if true, this has major implications for the industry's supply landscape and represents a rare "performance-cost" sweet spot in the open-source community
2026-08-25 ~ 2026-08-27 · 38 related posts
Primary sources
- [source] Zai Releases GLM-5.3-Flash Model on Hugging Face — zai-org · 2026-08-25
- Zhipu open-sources GLM-5.3-Flash: matches Claude Opus 4.8 at 1/40 the price — 智谱 · 2026-08-26
- [source] Zai Announces GLM-5.3-Flash Pricing: $0.15 per 1M Input Tokens — Zai_org · 2026-08-26
- [source] GLM-5.3-Flash Matches Claude Opus 4.8 on Code Bench — Zai_org · 2026-08-26
- GLM-5.3-Flash: Architecture Enhancements Boost Efficiency — Zai_org · 2026-08-26
- Zai Launches 320B Parameter GLM-5.3-Flash Model — TheZachMueller · 2026-08-26
- Zhipu releases GLM-5.3-Flash: 320B params, MIT open source — TheZachMueller · 2026-08-26
- Zhipu quietly releases glm-5.3-flash model — koltregaskes · 2026-08-26
- Zhipu's GLM-5.3-Flash Model Released on Hugging Face — coder543 · 2026-08-26
- Zhipu releases GLM-5.3-Flash: smaller size, Opus 4.8 level performance — airesearch12 · 2026-08-26
- GLM-5.3-Flash Uses Hybrid Attention to Cut Long-Context Costs — multimodalart · 2026-08-26
- GLM-5.3-Flash adopts sparse + linear attention hybrid architecture — multimodalart · 2026-08-26
- Zhipu Releases GLM-5.3-Flash: Frontier Intelligence, Flash Cost — dydynam · 2026-08-26
- Zhipu Launches GLM 5.3 Flash Vision: 320B Params at 1/40th of Opus 4.8's Price — oran_ge · 2026-08-26
- GLM-5.3 Flash: 320B total params, 18B active, deployed on Ascend — zephyr_z9 · 2026-08-26
- Zhipu Releases GLM 5.3: 100T Daily Compute on Domestic Chips, Costs Slashed by 90% — oran_ge · 2026-08-26
- GLM-5.3-Flash released with benchmark comparisons — elemental-mind · 2026-08-26
- Zhipu Releases GLM 5.3 Flash: 320B Params, 90% Price Cut — oran_ge · 2026-08-26
- Zai Releases GLM-5.3-Flash: 320B MoE with Hybrid Attention and 1M Context — TheZachMueller · 2026-08-26
- Zhipu Releases GLM-5.3-Flash: 320B MoE Open Source Under MIT — realsohamparekh · 2026-08-26
- GLM-5.3-Flash scores 57 on Artificial Analysis, hits Pareto frontier — ollama · 2026-08-26
- GLM-5.3-Flash matches Opus 4.8, 45x cheaper — Hesamation · 2026-08-26
- Zhipu GLM-5.3-Flash revealed: 320B params, MIT open source — op7418 · 2026-08-26
- Z.ai Releases GLM-5.3-Flash: Hybrid Sparse+Linear Attention Architecture — No_Afternoon_4260 · 2026-08-26
- New Qwen and GLM Open Models Benchmark Against Claude Opus — kimmonismus · 2026-08-26
- GLM-5.3-Flash Released: 320B MIT-Licensed Model Outperforms Predecessor — matei_zaharia · 2026-08-26
- Zhipu releases GLM-5.3-Flash: Frontier intelligence at flash cost — 1littlecoder · 2026-08-27
- Zhipu Open Sources GLM-5.3 Flash: Efficient 320B Model Ties Claude Opus — 量子位 · 2026-08-27
- GLM-5.3-Flash released: 320B MoE architecture, MIT licensed — TheZachMueller · 2026-08-27
- Zhipu GLM-5.3-Flash Pricing Revealed: $0.15 Input — bindureddy · 2026-08-27
- Zhipu Releases GLM-5.3-Flash: 320B MoE Model with MIT License — ramagetime · 2026-08-27
- GLM-5.3-Flash Review: 10% Cost, Pareto Frontier Performance — ArtificialAnlys · 2026-08-27
- GLM-5.3-Flash Released: 320B Total Params, Cost-Efficiency on Pareto Frontier — ArtificialAnlys · 2026-08-27
- GLM-5.3-Flash Released: Native Multimodal, 1M Context, 10x Cheaper — ying11231 · 2026-08-27
- Correction: GLM-5.3-Flash features a 1M context window — ArtificialAnlys · 2026-08-27
- GLM-5.3-Flash Released: 1M Context & Open Weights — qinzytech · 2026-08-27
- GLM-5.3-Flash Launches: 1M Context for $0.15 with Hybrid Attention — gharik · 2026-08-27
1 near-duplicate retellings: wandb