Zai Releases GLM-5.3-Flash: 320B MoE with Hybrid Attention and 1M Context
TheZachMueller · x · 2026-08-26
Zai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. It features a 320B total parameter / 18B active parameter MoE architecture, combining sparse and linear attention mechanisms with a 1M-token context window. The model is released under the MIT License and has day-0 support in vLLM, verified on both NVIDIA and AMD GPUs.
More from Models
- User fires Gemini after poor performance showcased in screenshot — dh7net · 2026-08-26
- 用户发现特定提示词可触发 Opus 4.6 及 GPT Image 1.5 — koltregaskes · 2026-08-26
- Qwen3.8-Flash-Next ranks top 8 in Code Arena, leading open weights in WebDev — arena · 2026-08-26
- GLM-5.3-Flash handles 100T tokens daily, running entirely on Chinese chips — airesearch12 · 2026-08-26
- Chinese lab ZAI cuts costs 10x by adopting peer innovations like DeepSeek — PAstynome · 2026-08-26
- Claude Fable 5 Starts Routing to Fable 5.1 for Some Users — legit_api · 2026-08-26