Chinese frontier models adopt linear attention; Zai releases MIT-licensed 320B GLM-5.3-Flash
ivan_bezdomny · x · 2026-08-26
Elie Bakouch observes that, except for DeepSeek and Kimi, all major Chinese frontier models have adopted linear attention and sparse attention designs with similar indexer/compression mechanisms. They also utilize advanced residuals (mHC, attention residual) and the Muon optimizer for efficiency. Concurrently, Zai released GLM-5.3-Flash (formerly Ox Alpha), a 320B-parameter (18B active) natively multimodal model with a 1M-token context window. It runs entirely on Chinese AI chips and is released under the MIT License.
More from Models
- Qwen3.8-Flash-Next released in 1-4bit GGUF formats — MaziyarPanahi · 2026-08-27
- Claude's "not just X, this is Y" tic likely comes from post-training, not web data — burkov · 2026-08-27
- Tiny 307M-parameter model outperforms 26x larger Qwen in embedding benchmarks — lateinteraction · 2026-08-27
- Inside GLM-5.3-flash: beats GLM-5.2 at 1/10 cost, active params halved to 18B — baseten · 2026-08-27
- Comparing tokens across different models is meaningless, metrics need refinement in the reasoning era — adamdangelo · 2026-08-27
- Yutori launches n2, a cost-effective computer-use model — DhruvBatra_ · 2026-08-27