MetaCanvas lands NeurIPS: MLLMs plan directly in diffusion latent space
mohitban47 · x · 2026-10-06
MetaCanvas is accepted to NeurIPS 2026. Instead of reducing MLLMs to global text encoders for diffusion models, it lets them reason and plan directly in spatial and spatiotemporal latent spaces.
- A set of learnable canvas tokens carries the MLLM's planning into the diffusion latent space via a lightweight connector; visualizations show they act as visual planning sketches guiding synthesis.
- Implemented on three diffusion backbones, evaluated across six tasks: text-to-image, text/image-to-video, image/video editing, and in-context video generation — each demanding precise layouts, robust attribute binding, and reasoning-intensive control.
- Consistently beats global-conditioning baselines, narrowing the gap between multimodal understanding and generation.
More from Multimodal
- Level-of-Token DiT Lets Pretrained Diffusion Models Use Arbitrary-Size Token Grids — GordonWetzstein · 2026-10-07
- Stanford's Level-of-Token Diffusion allocates fine tokens only where detail matters — GordonWetzstein · 2026-10-07
- Image-to-video model Prism with joint video-audio generation trends on Hugging Face, MIT licensed — FrancisRing · 2026-10-07
- iADD corrects DDPO theory: latter-timestep-only updates can harm diversity — Ashok Prasad Neupane · 2026-10-07
- Foresight lets streaming VLMs plan future perception without retraining, beating baselines by 9.5% — Ashok Prasad Neupane · 2026-10-07
- Qwen-Image-2.1 Pro lands on Runware: native 2K generation at $0.075 per image — aziz4ai · 2026-10-07