NVIDIA releases GLM-5.3-Flash NVFP4 quantized model on Hugging Face

TheZachMueller · x · 2026-09-10

NVIDIA has published an NVFP4 quantized version of GLM-5.3-Flash on Hugging Face, enabling low-precision inference on NVIDIA hardware. The release reuses Zhipu's reasoning-effort and tool-calling chat template; it's a quantization drop rather than a new model.

Original post →

More from Models

Models channel →