GLM-5.2 vision model baseten/GLM-5.2-Vision-NVFP4 trends on Hugging Face

baseten · hf · 2026-07-23

baseten/GLM-5.2-Vision-NVFP4 is trending on Hugging Face as an image-text-to-text model.

The entry points to a multimodal pipeline with sglang, safetensors, and basemodel:zai-org/GLM-5.2, highlighting a vision-language variant of the GLM family rather than a general text model.

Original post →

More from Multimodal

Multimodal channel →