Zhipu Releases GLM-5.3-Flash: 320B MoE Model with MIT License
ramagetime · x · 2026-08-27
Zhipu AI has released the GLM-5.3-Flash model, a 320B parameter sparse model (Active 18B), previously previewed as Ox Alpha and running entirely on Chinese AI chips.
Key highlights:
- High Performance at Low Cost: Priced at $0.15/1M input tokens, $0.50/1M output tokens, and $0.03 for cached input.
- Multimodal & Long Context: Natively supports 1M-token context window and multimodal capabilities.
- Fully Open Source: Released under the MIT license with weights and code available.
This release sets a new benchmark for cost-performance balance in open-weight models.
More from Models
- AI enters adolescence: Small models beating large ones in specific domains — DhruvBatra_ · 2026-08-27
- TerminalBench Deep Dive: GPT-5.6 Leads, Many Models Drop in Rank — abeirami · 2026-08-27
- Gemini 3.5 Transcribe Live Beats GPT Live Transcribe: 5.8% WER at 0.25s First Partial — ArtificialAnlys · 2026-08-27
- Qwen3.8-Flash-Next Runs on Dual DGX Sparks — NVIDIAAI · 2026-08-27
- New leader emerges on LLM leaderboard, claiming the crown — jonathan_wilke · 2026-08-27
- Yutori's Batra: most of the web will never get agent APIs — pixels in, clicks out — DhruvBatra_ · 2026-08-27