Zai Releases GLM-5.3-Flash: 320B MoE with Hybrid Attention and 1M Context

TheZachMueller · x · 2026-08-26

Zai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. It features a 320B total parameter / 18B active parameter MoE architecture, combining sparse and linear attention mechanisms with a 1M-token context window. The model is released under the MIT License and has day-0 support in vLLM, verified on both NVIDIA and AMD GPUs.

Related event: Zhipu Open-Sources GLM-5.3-Flash: 320B MoE Matching Opus 4.8 at 1/40 the Price(23 posts)→

Original post →

More from Models

Models channel →