NVIDIA and Waterloo release PixelUMM: an encoder-free multimodal model reading and writing raw pixels
CSProfKGD · x · 2026-10-02
NVIDIA and University of Waterloo introduce PixelUMM, an encoder-free unified multimodal model for image and video understanding and generation, with code and weights released today.
- Core idea: drop both VAEs and ViT vision encoders; a single decoder-only Transformer reads and writes raw pixels directly — images as 16×16 patches, videos as 4-frame tubes.
- Architecture: a Qwen3-8B backbone with separate understanding and generation experts shares one self-attention across text, clean pixels, and noisy pixels.
- Findings: eight experiment families (F1–F8) detail ablations. F3 shows linear pixel heads produce grid-aligned patch artifacts in low-texture regions like sky at high CFG (6); released checkpoints and demos still use the default linear heads, as switching to convolutional heads requires further training.
- Project page, paper, and model are all public — a notable entry in the encoder-free unified multimodal direction.
Related event: NVIDIA Releases PixelUMM, Encoder-Free Unified Multimodal Model(3 posts)→
More from Multimodal
- Indie dev generates an AI agent promo video with one prompt via Claude Code Skill — yihui_indie · 2026-10-02
- LeCun shares Photoroom's live try-on demo: clothes move with you, cutting returns — ylecun · 2026-10-02
- Claude Opus 5.5 directed a 5.5-minute AI short film from one prompt in ComfyUI — Cheap_Credit_3957 · 2026-10-02
- Skipping perfect pixels: MiniMax H3 renders 15-second shots in 6 minutes on an RTX 3060 — Support_Marmoset · 2026-10-02
- MachGen pushes MiniMax H3 past its 15s cap with 30-second continuous video — MiniMax_AI · 2026-10-02
- FLUX 3 Image launches with pixel-level control, multi-turn edits, 4K, and 10 reference images — Friendly-Fig-6015 · 2026-10-02