Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization

alejandroll10 · x · 2026-09-01

Z.ai has released the GLM-5.3-Flash model, featuring a hybrid sparse and linear attention architecture with 320B total parameters (18B active) and a 1M-token context window supporting text, image, and video input. Optimized for efficient coding and long-horizon agent tasks, it is now available via OrcaRouter API at $0.075/1M input tokens and $0.25/1M output tokens. Additionally, an uncensored NVFP4 version has been released to optimize inference performance on NVIDIA GPUs without safety filters.

Original post →

More from Models

Models channel →