GLM-5.3-Flash Released: Native Multimodal, 1M Context, 10x Cheaper
ying11231 · x · 2026-08-27
Zai officially released GLM-5.3-Flash (previously Ox Alpha). Features: First native multimodal model in the GLM-5 series with visual self-correction; supports 1M token context; 320B total (18B active) MoE architecture under MIT license. Performance: Outperforms GLM-5.2 at 1/10th the cost; hybrid sparse + linear attention enables efficient long-context. Capabilities: Extends beyond coding to professional workflows like slides, docs, spreadsheets, and finance research. Ecosystem: Day-0 support available via SGLang; weights, API, and tools are now open.
Related event: Zhipu open-sources GLM-5.3-Flash, a 320B MoE frontier model(38 posts)→
More from Models
- Rumor: Anthropic to launch Fable 5.1 before Astra is ready — haider1 · 2026-08-27
- Deepseek V4 Flash hits 420 tok/s in new community benchmark — HankYeomans · 2026-08-27
- Grok Bot gets more efficient with higher rate limits, users praise rapid improvement — XFreeze · 2026-08-27
- Users discuss stricter censorship in recent model updates — Connect-Cost-5504 · 2026-08-27
- Pokee-Isaac 28B Builds Playable Game in 5 Minutes with 10M Context — Kyrannio · 2026-08-27
- Goodfire AI Research: Efficiently Locating 'Forking Tokens' in LLMs — VoidAsuka · 2026-08-27