Stanford's Level-of-Token Diffusion cuts image and video generation cost with multiresolution tokens

GordonWetzstein · x · 2026-10-09

Stanford researchers with Google introduce Level-of-Token (LoT) Diffusion, replacing uniform token grids in diffusion transformers with a multiresolution layout: fine tokens where detail matters, coarse rectangular patches elsewhere. A patch-wise asymmetric flow parametrization plus multiresolution embeddings adapt pretrained models without losing full-resolution flow prediction at any denoising step. Layouts can come from semantic masks, bounding boxes, texture variance, depth-of-field cues, or agentic plans, yielding favorable quality-efficiency tradeoffs across image and video generation, with speedups set by the layout's token budget.

Original post →

More from Multimodal

Multimodal channel →