NVIDIA releases GLM-5.3-Flash NVFP4 quantized model on Hugging Face
TheZachMueller · x · 2026-09-10
NVIDIA has published an NVFP4 quantized version of GLM-5.3-Flash on Hugging Face, enabling low-precision inference on NVIDIA hardware. The release reuses Zhipu's reasoning-effort and tool-calling chat template; it's a quantization drop rather than a new model.
More from Models
- Codex CLI 0.154.0 ships GPT-6-Astra, experimental worktree support — github-actions[bot] · 2026-09-10
- OpenAI's new model, in training since Aug 28, reportedly beat GPT-6-Astra in just one week — alexcovo_eth · 2026-09-10
- Qwen3.8-2.4T-A95B open weights land on AWS: single 8×B300 node with vLLM — AWS ML Blog · 2026-09-10
- Users say Astra's $200 sub is no longer enough: multi-project work burns through quota in days — CtrlAltDwayne · 2026-09-10
- OpenAI launches GPT-6 Astra to power ChatGPT Work with desktop app control — OpenAI · 2026-09-10
- Intelligence Index v4.3: Claude Fable 5.1, Muse Spark 1.3 and GPT-6 Astra Reset the Cost-Efficiency Frontier — ArtificialAnlys · 2026-09-10