Level-of-Token DiT Lets Pretrained Diffusion Models Use Arbitrary-Size Token Grids
GordonWetzstein · x · 2026-10-07
Researchers introduce Level-of-Token DiT, a minimal modification to pretrained DiTs that supports non-uniform token layouts:
- The key idea generalizes the uniform token grid of pretrained DiTs into a "Level-of-Token" layout, where each token is a rectangle of any size and together they tile the whole image.
- Each token is projected from the latent patches it covers; the DiT is conditioned on token shapes, and extent-dependent heads restore the asymmetric velocity, which is converted into the full-rank velocity.
- They fine-tune both image (Flux.2) and video (Wan2.1) models with patch-wise asymmetric flow matching, preserving the pretrained generative prior.
Related event: Stanford's Level-of-Token Diffusion Speeds Up Generation Up to 4.6x(4 posts)→
More from Multimodal
- Voice Agents Live or Die on 'Sounding Right': Turbo Shifts Tone With User Emotion — SucceededMind · 2026-10-07
- Magnific Original Series The Chronicles of Bone drops Chapter Six, made entirely with AI tools — Kavanthekid · 2026-10-07
- Hedra Lands in ChatGPT: Attach One Product Photo, Get a Full Commercial Ad — henloitsjoyce · 2026-10-07
- Marc Andreessen boosts AI film contest SLOPTOBERFEST grand prize to $25,000 — zealcaiden · 2026-10-07
- Image generation pricing leak: $0.05 per 2K image, $0.076 per 4K — op7418 · 2026-10-07
- Live human votes plugged into Flow-GRPO to stop image models gaming reward models — lmoroney · 2026-10-07