Z.ai Releases GLM-5.3-Flash: Hybrid Sparse+Linear Attention Architecture

No_Afternoon_4260 · reddit · 2026-08-26

Z.ai released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the first open-weight release of the glm5next architecture. Pitched as outperforming GLM-5.2 at one-tenth the price and approaching Claude Opus 4.8 on coding and agentic benchmarks.

Key Highlights:

Specs: 320B total params (18B active), MoE (288 routed, 8 active), 1M context window, MIT License.

Related event: Zhipu's GLM-5.3-Flash: 320B params, coding matches Claude Opus 4.8(24 posts)→

Original post →

More from Models

Models channel →