GLM-5.2 vision model baseten/GLM-5.2-Vision-NVFP4 trends on Hugging Face
baseten · hf · 2026-07-23
baseten/GLM-5.2-Vision-NVFP4 is trending on Hugging Face as an image-text-to-text model.
The entry points to a multimodal pipeline with sglang, safetensors, and basemodel:zai-org/GLM-5.2, highlighting a vision-language variant of the GLM family rather than a general text model.
More from Multimodal
- FLUX 3 can invent its own camera cuts from a one-line image prompt — venturetwins · 2026-07-23
- Self-Flow speeds multimodal model convergence by up to 2.8×, paper says — hila_chefer · 2026-07-23
- Viewer gifts now trigger real-time AI video effects in a live-streaming demo — ming_calligraphy · 2026-07-23
- Controlled study finds training data quality is decisive for text-to-video models — Amber Yijia Zheng · 2026-07-23
- FLUX 3 is announced, but its capabilities will roll out over weeks and months — Angaisb_ · 2026-07-23
- Black Forest Labs positions FLUX 3 as a multimodal backbone for visual intelligence — stephen370 · 2026-07-23