Zhipu's GLM-5.3-Flash: 320B params, coding matches Claude Opus 4.8
Zhipu officially released and open-sourced GLM-5.3-Flash on 08-26 — the mysterious model "Ox Alpha" that had stirred buzz on OpenRouter. Delivering frontier performance at a fraction of flagship pricing, it became the day's most-watched model launch.
Confirmed
- Architecture: a MoE with 320B total and 18B active parameters; the first natively multimodal model in the GLM-5 series, supporting a 1 million token context window. It uses a hybrid of sparse and linear attention, significantly cutting long-context serving costs while preserving accurate long-context capability (@Zaiorg, @multimodalart).
- Performance: scores 57 on the Artificial Analysis Intelligence Index, on par with Claude Opus 4.8 and on the Pareto frontier; on Code Bench real-world coding benchmarks it significantly outperforms GLM-5.2 across all compute tiers and matches Opus 4.8 (@Zaiorg, @op7418).
- Pricing: standard API at $0.15 per million tokens input, $0.50 output, $0.03 cached input; @Hesamation estimates roughly $0.045 per task — 15x cheaper than GLM-5.3 and about 45x cheaper than Opus 4.8.
- Open source: released under MIT; FP8 weights take about 328GB, currently deployed on Ascend hardware, not yet NVFP4 (@zephyrz9).
- Compute: @orange relays official claims that the entire 100T+ per-day global supply is backed by domestic Chinese chips, with 3x end-to-end serving performance gains.
- Ecosystem: OpenRouter and other platforms are on board; the model had already consumed 23.2T tokens on OpenRouter, ending DeepSeek's reign atop consumption charts (@op7418); vLLM support is available.
Why it matters
- Achieving an intelligence index on par with Claude Opus 4.8 at roughly 1/40 the price reshapes the performance-per-dollar Pareto frontier and puts direct pressure on closed-source flagship API pricing.
- Open-sourcing a 320B MoE multimodal model under MIT, backed entirely by domestic Chinese chips, is seen as a signal that could reshape the industry.
- Architecturally, the hybrid attention approach lowers long-context serving costs and — alongside contemporaries like Qwen3.8-Flash-Next — represents a new wave of efficient inference models.
2026-08-26 ~ 2026-08-26 · 23 related posts
- Episode 1: Mystery model Ox Alpha stuns community; fingerprints point to Zhipu GLM-5.3(2026-08-21, 58 posts)
- Episode 2: OpenRouter's Mystery 'Ox Alpha' Model Identified as GLM(2026-08-21, 3 posts)
- Episode 3: Cola Launches Ox Model Claiming Terra-Level Capability at Flash Speed(2026-08-21, 2 posts)
- Episode 4: Mysterious Model Ox Alpha Gives Away 100T Tokens Daily, Sparking Power Questions(2026-08-21, 2 posts)
- Episode 5: Ox-Alpha Underperforms in Tests, Called Marketing Stunt(2026-08-22, 3 posts)
- Episode 6: Ox Alpha scores ~63% on full DeepSWE; earlier 80% was subset artifact(2026-08-22, 6 posts)
- Episode 7: Forensics Point to Zhipu's Unreleased GLM Behind Mystery Model Ox Alpha(2026-08-22, 31 posts)
- Episode 8: Ox Alpha Generates 64K-Token 3D Scene in One Prompt(2026-08-22, 2 posts)
- Episode 9: Mystery model Ox Alpha draws mixed reviews from early testers(2026-08-22, 6 posts)
- Episode 10: Mystery Model Ox Alpha Launches Free, Identity Unknown(2026-08-23, 3 posts)
- Episode 11: Mysterious Model Ox Alpha Debuts on OpenRouter, Sparking Speculation(2026-08-24, 5 posts)
- Episode 12: Leak: Ox Alpha Is GLM-5.3 Flash, 320B MoE at One-Tenth the Price(2026-08-25, 4 posts)
- Episode 13: Mystery chart-topper Ox Alpha confirmed as Zhipu GLM-5.3-Flash, weights open-sourced tonight(2026-08-25, 16 posts)
- Episode 14: Stealth Model Ox Alpha Proves Mediocre at Coding, Trails Grok 4.6(2026-08-25, 3 posts)
- Episode 15: Mystery Model Ox Alpha Burns 42T Tokens in Six Days(2026-08-26, 3 posts)
- Episode 16: Zhipu's GLM-5.3-Flash: 320B params, coding matches Claude Opus 4.8(2026-08-26, 23 posts)
Primary sources
- [source] Zhipu open-sources GLM-5.3-Flash: matches Claude Opus 4.8 at 1/40 the price — 智谱 · 2026-08-26
- [source] Zai Announces GLM-5.3-Flash Pricing: $0.15 per 1M Input Tokens — Zai_org · 2026-08-26
- [source] GLM-5.3-Flash Matches Claude Opus 4.8 on Code Bench — Zai_org · 2026-08-26
- GLM-5.3-Flash: Architecture Enhancements Boost Efficiency — Zai_org · 2026-08-26
- Zai Launches 320B Parameter GLM-5.3-Flash Model — TheZachMueller · 2026-08-26
- Zhipu releases GLM-5.3-Flash: 320B params, MIT open source — TheZachMueller · 2026-08-26
- Zhipu quietly releases glm-5.3-flash model — koltregaskes · 2026-08-26
- Zhipu releases GLM-5.3-Flash: smaller size, Opus 4.8 level performance — airesearch12 · 2026-08-26
- GLM-5.3-Flash Uses Hybrid Attention to Cut Long-Context Costs — multimodalart · 2026-08-26
- GLM-5.3-Flash adopts sparse + linear attention hybrid architecture — multimodalart · 2026-08-26
- Zhipu Releases GLM-5.3-Flash: Frontier Intelligence, Flash Cost — dydynam · 2026-08-26
- Zhipu Launches GLM 5.3 Flash Vision: 320B Params at 1/40th of Opus 4.8's Price — oran_ge · 2026-08-26
- GLM-5.3 Flash: 320B total params, 18B active, deployed on Ascend — zephyr_z9 · 2026-08-26
- Zhipu Releases GLM 5.3: 100T Daily Compute on Domestic Chips, Costs Slashed by 90% — oran_ge · 2026-08-26
- GLM-5.3-Flash released with benchmark comparisons — elemental-mind · 2026-08-26
- Zhipu Releases GLM 5.3 Flash: 320B Params, 90% Price Cut — oran_ge · 2026-08-26
- Zai Releases GLM-5.3-Flash: 320B MoE with Hybrid Attention and 1M Context — TheZachMueller · 2026-08-26
- Zhipu Releases GLM-5.3-Flash: 320B MoE Open Source Under MIT — realsohamparekh · 2026-08-26
- GLM-5.3-Flash scores 57 on Artificial Analysis, hits Pareto frontier — ollama · 2026-08-26
- GLM-5.3-Flash matches Opus 4.8, 45x cheaper — Hesamation · 2026-08-26
- Zhipu GLM-5.3-Flash revealed: 320B params, MIT open source — op7418 · 2026-08-26
- Z.ai Releases GLM-5.3-Flash: Hybrid Sparse+Linear Attention Architecture — No_Afternoon_4260 · 2026-08-26
- New Qwen and GLM Open Models Benchmark Against Claude Opus — kimmonismus · 2026-08-26