Diffusion Transformers' contextual tokens already encode the image before it's visible

kwangmoo_yi · x · 2026-10-07

A Tel Aviv University and Cornell paper trains a lightweight bottleneck network to map MM-DiT contextual tokens into a frozen LLM's input space, revealing that these tokens encode rich semantics about the emerging image even at early denoising stages. Supervising the tokens with this readout improves generation quality and distributional coverage.

Related event: Study: Diffusion Transformers Encode Image Content Early in Contextual Tokens(2 posts)→

Original post →

More from Multimodal

Multimodal channel →