NVIDIA's PixelUMM: encoder-free unified image and video understanding and generation
nvidia · hf · 2026-10-02
NVIDIA introduces PixelUMM, an encoder-free unified multimodal model that performs image and video understanding and generation directly in pixel space.
- Motivation: existing UMMs use separate visual representations for understanding vs. generation, inflating visual context length and complicating integration with vision-language pretraining; extending pixel-space modeling from images to video is non-trivial due to differing temporal representations
- Method: images as spatial patches, videos as spatiotemporal tubelets, connected via single-layer linear projections to a shared backbone; a Mixture-of-Transformers architecture combines shared attention with task-specific parameters, extending clean-pixel prediction to video generation and jointly supporting autoregressive text prediction and pixel-space flow matching
- Results: competitive performance across image and video understanding and generation, plus ablations on decoder design and spatiotemporal patch size
Related event: NVIDIA Releases PixelUMM, Encoder-Free Unified Multimodal Model(3 posts)→
More from Multimodal
- MachGen pushes MiniMax H3 past its 15s cap with 30-second continuous video — MiniMax_AI · 2026-10-02
- NVIDIA and Waterloo release PixelUMM: an encoder-free multimodal model reading and writing raw pixels — CSProfKGD · 2026-10-02
- One Prompt Rewrite Turned AI Video Characters From 'Moving' Into Actually 'Acting' — FinanceYF5 · 2026-10-02
- Fable 5.5 outputs so good that even 'slop creators' are making insane videos — inductionheads · 2026-10-02
- PixVerse Shares Full Prompt for GPT Image 2.5 Architectural Icons Posters — aziz4ai · 2026-10-02
- One word change, four styles: HeyGen video generation demo wows users — aziz4ai · 2026-10-02