NVIDIA Releases PixelUMM, Encoder-Free Unified Multimodal Model
NVIDIA and the University of Waterloo released PixelUMM, an open-source encoder-free multimodal model built on Qwen3-8B that handles image and video understanding and generation directly in pixel space.
2026-10-02 ~ 2026-10-02 · 3 related posts
- NVIDIA releases PixelUMM on Hugging Face, a Qwen3-8B-based image-to-text model — nvidia · 2026-10-02
- NVIDIA's PixelUMM: encoder-free unified image and video understanding and generation — nvidia · 2026-10-02
- NVIDIA and Waterloo release PixelUMM: an encoder-free multimodal model reading and writing raw pixels — CSProfKGD · 2026-10-02