Zhipu & Alibaba release new Flash models, redefining cost-performance

APPSO · wechat · 2026-08-28

Zhipu open-sourced GLM-5.3-Flash (formerly Ox-Alpha), and Alibaba released Qwen3.8-Flash (featuring the new Qwen4 architecture). Both retain million-token context and multimodal capabilities at significantly lower costs: GLM-5.3-Flash is priced at 1/20th of GLM-5.3, and Qwen3.8-Flash at 1/12th of Qwen3.8-Max. Tests show GLM-5.3-Flash performs similarly to flagship models in coding and multimodal tasks, powered entirely by domestic chips. The article argues that Flash models are evolving from stripped-down versions to efficient architectures (like MoE), shifting industry competition to a "high intelligence x low cost" baseline.

Original post →

More from Venture

Venture channel →