Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization
alejandroll10 · x · 2026-09-01
Z.ai has released the GLM-5.3-Flash model, featuring a hybrid sparse and linear attention architecture with 320B total parameters (18B active) and a 1M-token context window supporting text, image, and video input. Optimized for efficient coding and long-horizon agent tasks, it is now available via OrcaRouter API at $0.075/1M input tokens and $0.25/1M output tokens. Additionally, an uncensored NVFP4 version has been released to optimize inference performance on NVIDIA GPUs without safety filters.
More from Models
- Focus on specific tasks, not the best model, as selection logic evolves — aftahi_ai · 2026-09-01
- User Rants on GPT-5.6 Hallucinations and Coding Limits, Hopes for GPT-6 Fix — Prestigiouspite · 2026-09-01
- Rumor: GPT-6 'Astra' nears human-level computer use — jYtanYj · 2026-09-01
- Open Source Models Shift to Revenue Sharing and Licensing — zephyr_z9 · 2026-09-01
- MiniMax Hailuo H3 Max is fast enough to power a playable AI open-world RPG — mtizard · 2026-09-01
- September may see biggest model release wave; Anthropic, OpenAI updates imminent — ChrisUniverse · 2026-09-01