Seedream team unveils VoT: visual reasoning branch plans images before diffusion renders pixels

JingxiangSun42 · x · 2026-09-13

The ByteDance Seedream team introduces VoT: Vision-of-Thought for Unified Multimodal Representation Alignment, extending MoT into three branches — Text, VoT, and Diffusion — with key-value pairs from all branches concatenated for unified attention.

For text-to-image and image-to-image generation, the VoT branch between the VLM and DiT first predicts discrete visual tokens as semantic plans ("visual thinking before rendering pixels"), then the diffusion branch jointly attends to the text and VoT key-value pairs to synthesize the final image.

Related event: ByteDance Seedream Team Introduces VoT: Thinking Before Rendering(3 posts)→

Original post →

More from Multimodal

Multimodal channel →