Z.ai Releases Open-Weight GLM-5.3 with Day-One Adoption Across Platforms
Z.ai released GLM-5.3 on August 29, its strongest open-weight model for agentic coding and cyber defense, available for download, deployment and commercial use. Built on the same 743B-A40B (744B in some Baseten posts) MoE architecture as GLM-5.2 with extended post-training on realistic tasks such as full codebases, it is stronger at coding and long-horizon agent tasks and supports a 256k context (Baseten posts cite 1 million tokens). The team ran an extra two-week safety evaluation due to its strong cyber capabilities, and the model saw day-one adoption across major inference, fine-tuning and application platforms.
Confirmed
- License switched from MIT to GLM-MIT: MaaS providers with over $10 billion annual revenue face mandatory safety review; SGLang announced day-one serving support
- Performance: Terminal Bench 3.0 rose from 4.6 to 28.3 on Cloudflare Workers AI (open-source SOTA); Together AI says it nearly matches Fable 5 at far lower per-task cost; Perplexity reports it beats GLM 5.2 on its WANDR evidence-research benchmark; Tinker posts cite strong results on Terminal-Bench 3.0 and DeepSWE 1.1
- Pricing: 100 tps on engy.ai at $0.98/M input and $3.08/M output tokens (about 30% below peers); an 8-bit quantized version is offered at $0.40/M input, $0.06/M cache, $1.40/M output
- Day-one platforms: Cloudflare Workers AI, Baseten (US-only, ZDR), Databricks (Day 0, own secure GPUs, Unity Gateway agent routing), Together AI (Verified Inference), Tinker (first GLM model, fine-tuning), Modal, Perplexity Computer, and Ollama cloud (GLM 5.3 and Flash, private deployment, EU/US hosting, zero retention, Claude Desktop integration)
- Lightweight GLM 5.3 Flash (formerly Ox Alpha): 320B MoE with 18B active parameters, live on Phala with TDX-certified GPU TEE private inference
Unconfirmed
- License reports conflict: GLM-MIT vs MIT per Baseten posts; official repo governs
- Parameter counts (743B vs 744B) and context length (256k vs 1M tokens) vary across posts
Why it matters
Widely called the smartest open-weight model, GLM-5.3 sets new open-source records on agent coding benchmarks at sharply lower prices, with a Flash variant and TEE private inference — likely intensifying competition between open-weight models and closed APIs in coding and agent workloads.
2026-08-29 ~ 2026-08-30 · 19 related posts
- Episode 1: Zhipu GLM-5.3 Spotted in Codebase, Post-Training Targets K3 and GPT-5.6(2026-08-03, 5 posts)
- Episode 2: Zhipu AI Releases GLM-5.3: Same Base Model, Post-Training Drives a Big Leap(2026-08-14, 51 posts)
- Episode 3: AI Weekly: Grok, Gemini Updates, and Anthropic's Secret Model(2026-08-14, 2 posts)
- Episode 4: Alibaba, Zhipu and DeepSeek Unveil New Models on the Same Day(2026-08-15, 3 posts)
- Episode 5: Zhipu Launches GLM-5.3 API with 50% Coding Boost(2026-08-19, 4 posts)
- Episode 6: GLM-5.3 Scores 60 on AA Intelligence Index, Tops Agentic Index(2026-08-19, 12 posts)
- Episode 7: Zhipu GLM-5.3 Tops DeepSWE at a Fraction of the Cost(2026-08-20, 4 posts)
- Episode 8: Zhipu Open-Sources GLM-5.3-Flash: 320B MoE Multimodal at 1/10th the Cost(2026-08-25, 48 posts)
- Episode 9: GLM-5.3 Flash Benchmarks Hit GPT-5.6 Territory at Fractional Cost(2026-08-27, 7 posts)
- Episode 10: GLM-5.3-Flash Ranks Third Among Open Models at Half of Qwen's Cost(2026-08-27, 2 posts)
- Episode 11: GLM 5.3 Flash Gets 90% Discount Through September(2026-08-27, 2 posts)
- Episode 12: Zai Releases GLM-5.3 Weights with Tiered Access Citing Cybersecurity Capabilities(2026-08-27, 11 posts)
- Episode 13: GLM-5.3 Flash Benchmarks: Near-Frontier Quality at a Fraction of the Cost(2026-08-27, 12 posts)
- Episode 14: Z.ai Releases Open-Weight GLM-5.3 with Day-One Adoption Across Platforms(2026-08-29, 19 posts)
- Episode 15: Zhipu's GLM-5.3 Launches Exclusively on Baseten with Big Coding Gains(2026-09-05, 2 posts)
Primary sources
- GLM-5.3 launches on Baseten with 743B params and MIT license — baseten · 2026-08-29
- GLM 5.3 Flash launches on Phala with GPU TEE private inference — bgmshana · 2026-08-29
- GLM-5.3 on Cloudflare Workers AI: Doubles Long-Horizon Coding Performance — michellechen · 2026-08-29
- [source] GLM 5.3 Available in Perplexity Computer, Beats GLM 5.2 on WANDR Benchmark — perplexity_ai · 2026-08-29
- [source] GLM-5.3 launches on Baseten with major coding gains and 1M context — baseten · 2026-08-29
- GLM-5.3 Launches on Tinker with 256k Context — simonguozirui · 2026-08-29
- Zai's flagship model GLM 5.3 is now available on Modal — AAAzzam · 2026-08-29
- GLM-5.3 fine-tuning now available on Tinker platform — Zai_org · 2026-08-29
- [source] Zhipu GLM 5.3 and Flash Now Available on Ollama Cloud — ollama · 2026-08-29
- Databricks Launches GLM 5.3 with Agent Routing via Unity Gateway — pwendell · 2026-08-29
- GLM-5.3 Launches on Together AI with High Reasoning Efficiency and Lower Cost — togethercompute · 2026-08-29
- Zhipu GLM 5.3 and Flash models now available on Ollama cloud — johnseach · 2026-08-29
- Zai releases GLM-5.3 as open-weight for agentic coding and cyber defense — alexcovo_eth · 2026-08-29
- GLM-5.3 adopts GLM-MIT license, requiring security review for hyperscalers — AccBalanced · 2026-08-29
- GLM-5.3 Released for Agentic Coding at 30% Lower Cost — markjeffrey · 2026-08-30
- GLM-5.3 launched at $0.40/m input tokens — alejandroll10 · 2026-08-30