NVIDIA releases PixelUMM on Hugging Face, a Qwen3-8B-based image-to-text model

nvidia · hf · 2026-10-02

NVIDIA has released PixelUMM on Hugging Face, an image-to-text multimodal understanding model fine-tuned from Qwen/Qwen3-8B. It supports image understanding, video understanding, and video-text-to-text tasks, licensed under a custom 'other' license (region: us).

Related event: NVIDIA Releases PixelUMM, Encoder-Free Unified Multimodal Model(3 posts)→

Original post →

More from Models

Models channel →