Nvidia Releases GLM-5.2 NVFP4 Quantized Version

nvidia · hf · 2026-07-05

Nvidia released nvidia/GLM-5.2-NVFP4 on Hugging Face. Using its Model Optimizer/ModelOpt, Nvidia quantized Zhipu's GLM-5.2 (MoE) into NVFP4 4-bit precision, reducing memory footprint and inference costs while showcasing the adaptation and deployment of the NVFP4 format for domestic large models.

Original post →

More from Models

Models channel →