ZAI Launches GLM-5.3 Flash: 320B Native Multimodal Model
togethercompute · x · 2026-08-28
ZAI released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. It features 320B total parameters with 18B active, a 1M context window, and a hybrid attention architecture. On DeepSWE, it nearly matches Luna's performance while completing more than twice the work for the same budget. The model integrates vision for inspecting outputs and reading professional documents, available now on Together AI.
Related event: Zai Open-Sources GLM-5.3-Flash: 320B-Parameter Native Multimodal MoE Model(5 posts)→
More from Models
- GLM-5.3 Flash High Reasoning Live on HF Providers; Devs Call It Opus 4.8-Class — _akhaliq · 2026-08-28
- GLM-5.3 Open Weights: SGLang Reports Over 500 tok/s on Agentic Workloads — BanghuaZ · 2026-08-28
- GLM-5.3 released as open-weight model for agentic coding — shaunralston · 2026-08-28
- Gemini 1.5 Flash inference speeds may exceed 300 tok/sec — Sentdex · 2026-08-28
- Google open-sources TimesFM for zero-shot forecasting of trends — mdancho84 · 2026-08-28
- Hands-on: Tencent Hy4 preview shows significant research gains over Hy3 — ShunyuYao12 · 2026-08-28