GLM-5.3-Flash Released: 320B MIT-Licensed Model Outperforms Predecessor
matei_zaharia · x · 2026-08-26
Zai released GLM-5.3-Flash (Ox Alpha), a 320B-A18B model that is less than half the size of GLM-5.2 yet beats it across every benchmark. The model features native multimodality, a 1M-token context window, and is released under the MIT License. Notably, it runs entirely on Chinese AI chips. Databricks has announced plans to integrate GLM-5.3-Flash for customers quickly.
More from Models
- Qwen3.8-Flash-Next released in 1-4bit GGUF formats — MaziyarPanahi · 2026-08-27
- Claude's "not just X, this is Y" tic likely comes from post-training, not web data — burkov · 2026-08-27
- Tiny 307M-parameter model outperforms 26x larger Qwen in embedding benchmarks — lateinteraction · 2026-08-27
- Inside GLM-5.3-flash: beats GLM-5.2 at 1/10 cost, active params halved to 18B — baseten · 2026-08-27
- Comparing tokens across different models is meaningless, metrics need refinement in the reasoning era — adamdangelo · 2026-08-27
- Yutori launches n2, a cost-effective computer-use model — DhruvBatra_ · 2026-08-27