Z.AI Unveils GLM-5.3-Flash: 320B Native Multimodal Model with 90% Cost Cut
baseten · x · 2026-08-27
Z.AI has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series with 320B total parameters (18B active). It approaches Claude Opus 4.8 in coding and agentic benchmarks, offers 90% lower costs than GLM-5.2, and supports 1M token context with native vision. Baseten has launched the model API on day one.
More from Models
- Users Report GPT-5.6 Quality Drop Ahead of Potential Astra Release — Realistic_Stomach848 · 2026-08-27
- Claude tightens undocumented generation restrictions, impacting design outputs — max_paperclips · 2026-08-27
- Peter Yang tests ChatGPT, Claude, Grok and Gemini across 10 use cases — nickbaumann_ · 2026-08-27
- MiniMax H3 at Ray Summit: Showcasing 33B Open-Weight Audio-Video Model — MiniMax_AI · 2026-08-27
- MiniMax H3 Max tops video leaderboards via fal's post-training — ArtificialAnlys · 2026-08-27
- GLM-5.3 Flash Review: GPT-5.6 Level Performance at Ultra-Low Cost — zainhas · 2026-08-27