NVIDIA and Waterloo release PixelUMM: an encoder-free multimodal model reading and writing raw pixels

CSProfKGD · x · 2026-10-02

NVIDIA and University of Waterloo introduce PixelUMM, an encoder-free unified multimodal model for image and video understanding and generation, with code and weights released today.

Related event: NVIDIA Releases PixelUMM, Encoder-Free Unified Multimodal Model(3 posts)→

Original post →

More from Multimodal

Multimodal channel →