NVIDIA ships NVFP4-quantized GLM-5.3-Flash, trending on Hugging Face

nvidia · hf · 2026-09-16

NVIDIA released an NVFP4 (FP4) quantized build of GLM-5.3-Flash on Hugging Face, where it is trending. The image-text-to-text model was produced with NVIDIA's Model Optimizer (ModelOpt) toolkit and ships as safetensors, enabling lower-memory FP4 inference on NVIDIA hardware.

Original post →

More from Infra

Infra channel →