Chinese frontier models adopt linear attention; Zai releases MIT-licensed 320B GLM-5.3-Flash

ivan_bezdomny · x · 2026-08-26

Elie Bakouch observes that, except for DeepSeek and Kimi, all major Chinese frontier models have adopted linear attention and sparse attention designs with similar indexer/compression mechanisms. They also utilize advanced residuals (mHC, attention residual) and the Muon optimizer for efficiency. Concurrently, Zai released GLM-5.3-Flash (formerly Ox Alpha), a 320B-parameter (18B active) natively multimodal model with a 1M-token context window. It runs entirely on Chinese AI chips and is released under the MIT License.

Original post →

More from Models

Models channel →