GLM-5.3-Flash released: 320B MoE architecture, MIT licensed

TheZachMueller · x · 2026-08-27

GLM-5.3-Flash, previously known as Ox Alpha, is released. It features a 320B-A18B sparse MoE backbone, 1M-token context, native multimodality, and an MIT license. The architecture integrates Kimi Linear hybrid attention, DeepSeek Sparse Attention, and an mHC residual path, running entirely on Chinese AI chips.

Related event: Zhipu Open-Sources GLM-5.3-Flash: 320B MoE with 1M Context at One-Tenth the Cost(33 posts)→

Original post →

More from Models

Models channel →