Zhipu & Alibaba release new Flash models, redefining cost-performance
APPSO · wechat · 2026-08-28
Zhipu open-sourced GLM-5.3-Flash (formerly Ox-Alpha), and Alibaba released Qwen3.8-Flash (featuring the new Qwen4 architecture). Both retain million-token context and multimodal capabilities at significantly lower costs: GLM-5.3-Flash is priced at 1/20th of GLM-5.3, and Qwen3.8-Flash at 1/12th of Qwen3.8-Max. Tests show GLM-5.3-Flash performs similarly to flagship models in coding and multimodal tasks, powered entirely by domestic chips. The article argues that Flash models are evolving from stripped-down versions to efficient architectures (like MoE), shifting industry competition to a "high intelligence x low cost" baseline.
More from Venture
- Polymarket predicts 12% chance AI bubble bursts by end of 2026 — Polymarket · 2026-08-28
- Electric AI relaunches as AI-native with embedded GTM strategy — DenehyXXL · 2026-08-28
- GPU Gold Rush: Are Tech Companies Creating a Hype-Driven Bubble? — DavidLinthicum · 2026-08-28
- French dev raises over $3,000 in 24h auctioning MacBook lid ad space — Polymarket · 2026-08-28
- Speculation: OpenAI might acquire Salesforce for Slack to own context — GregKamradt · 2026-08-28
- Analyst: Semi Demand Exceeds Capacity by 15-20%, NVDA Supply Constrained — BenBajarin · 2026-08-28