Zhipu open-sources GLM-5.3-Flash: matches Claude Opus 4.8 at 1/40 the price
智谱 · wechat · 2026-08-26
Zhipu has released and open-sourced GLM-5.3-Flash (320B-A18B), the first natively multimodal model in the GLM-5 family. It scores 57 on the Artificial Analysis Intelligence Index, entering the global frontier band and matching Claude Opus 4.8, with comparable coding performance on Z.ai CodeBench — at 1/40 of Opus pricing (1/20 of GLM-5.3 during a limited discount). It beats the twice-as-large GLM-5.2 across benchmarks.
Key points:
- Architecture: first open-source frontier model combining sparse and linear attention, plus Manifold-Constrained Hyper-Connections (mHC) for better scaling; trained on 30T tokens of multimodal data. Attention compute and KV cache drop 3.01x and 4.44x vs GLM-5.3.
- Visual coding: natively integrated vision lets the model inspect and iterate on its own outputs; it autonomously built a 400㎡ chef's kitchen in Blender over 16 hours, and Browser/Computer Use Agents verify rendering and interaction.
- Domestic chips: all traffic was served on Chinese accelerator clusters; a custom SGLang-based engine with W8A8 quantization, mixed cache quantization and an Encode–Prefill–Decode disaggregated architecture delivered 3x end-to-end speedup, reaching per-token costs on par with mainstream NVIDIA GPUs.
- Availability: API, GLM Coding Plan (10,000 daily trial cards) and open weights are live. Pre-release, it was tested anonymously as Ox-Alpha on OpenCode and OpenRouter, becoming the most popular model of the week.
More from Infra
- ZAI acquires Zhongke Jiahe to boost efficiency on Huawei Ascend and Chinese accelerators — himanshustwts · 2026-08-26
- Autonomous Launches Personal AI Datacenter Hardware — dee_hw · 2026-08-26
- Opus-4.8 tier models trained and deployed on Chinese domestic AI chips — chris_j_paxton · 2026-08-26
- GLM-5.3-Flash handles 100T tokens daily, running entirely on Chinese chips — airesearch12 · 2026-08-26
- Chinese lab ZAI cuts costs 10x by adopting peer innovations like DeepSeek — PAstynome · 2026-08-26
- TAO Introduces Machine-Native Infrastructure for Autonomous Agents — bittingthembits · 2026-08-26