GLM-5.3-Flash released: 320B MoE architecture, MIT licensed
TheZachMueller · x · 2026-08-27
GLM-5.3-Flash, previously known as Ox Alpha, is released. It features a 320B-A18B sparse MoE backbone, 1M-token context, native multimodality, and an MIT license. The architecture integrates Kimi Linear hybrid attention, DeepSeek Sparse Attention, and an mHC residual path, running entirely on Chinese AI chips.
More from Models
- Gemini 3.5 Transcribe launches; Enterprise Agent Platform enters public preview — Saboo_Shubham_ · 2026-08-27
- Benchmarking Qwen3.8 27B Quantizations: 4-bit Holds Up, 1-bit Collapses — pmigdal · 2026-08-27
- GLM-5.3-Flash Review: 10% Cost, Pareto Frontier Performance — ArtificialAnlys · 2026-08-27
- Google announces pricing details for Gemini 3.7 Flash — OfficialLoganK · 2026-08-27
- Unsloth releases GGUF quantization of GLM-5.3-Flash model — unsloth · 2026-08-27
- Will N-Gram tables revolutionize the local AI race? — AcreMakeover · 2026-08-27