Ollama adds Z.ai's GLM-5.3-Flash: 18B active params, 1M context, near Opus 4.8
ollama · x · 2026-08-28
Ollama's cloud now hosts Z.ai's GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total and only 18B active parameters, it reportedly beats GLM-5.2 at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic benchmarks. Before release it was tested anonymously as ox-alpha, becoming the most popular model of the week. Features a 1M context window, vision, tools, and thinking; plug-and-play with Claude Code, OpenCode and other agent apps.
Related event: Zai Open-Sources GLM-5.3-Flash: 320B-Parameter Native Multimodal MoE Model(5 posts)→
More from Models
- Zai Releases Open-Weight GLM-5.3 Coding Model Amid Safety Criticism — dhadfieldmenell · 2026-08-28
- Claude vs Gemini: Both Hit 98.67% Semantic Pass Rate but Fail Differently — Plastic-Cell-4497 · 2026-08-28
- GLM-5.3 Flash High Reasoning Live on HF Providers; Devs Call It Opus 4.8-Class — _akhaliq · 2026-08-28
- Local Inference Tool ds4 Adds Support for GLM 5.3 Flash — lakySK · 2026-08-28
- Gemini 1.5 Flash inference speeds may exceed 300 tok/sec — Sentdex · 2026-08-28
- Google open-sources TimesFM for zero-shot forecasting of trends — mdancho84 · 2026-08-28