Stanford's Level-of-Token Diffusion cuts image and video generation cost with multiresolution tokens
GordonWetzstein · x · 2026-10-09
Stanford researchers with Google introduce Level-of-Token (LoT) Diffusion, replacing uniform token grids in diffusion transformers with a multiresolution layout: fine tokens where detail matters, coarse rectangular patches elsewhere. A patch-wise asymmetric flow parametrization plus multiresolution embeddings adapt pretrained models without losing full-resolution flow prediction at any denoising step. Layouts can come from semantic masks, bounding boxes, texture variance, depth-of-field cues, or agentic plans, yielding favorable quality-efficiency tradeoffs across image and video generation, with speedups set by the layout's token budget.
More from Multimodal
- ComfyUI app ANIMA goes viral: Qwen 3.5 'hallucinates' popular photos, 300K views — Ok_Contribution8157 · 2026-10-09
- NAMVIS (NeurIPS 2026): next-scale autoregression beats diffusion for novel-view synthesis, 3x faster — RexDouglass · 2026-10-09
- fal Engineer on Video Speed: H3 Max Renders 15 Seconds of Video in 5 — OdinLovis · 2026-10-09
- Odyssey launches Odyssey-3, claims SOTA world model on Physics-IQ benchmark — Scobleizer · 2026-10-09
- One prompt, one full movie: dev shows multi-director agent pipeline for AI films — Exciting-Income-5840 · 2026-10-09
- Nano Banana 2.1 becomes Google's best image editing model across all 7 edit actions — ArtificialAnlys · 2026-10-09