Marigold V2 arrives as Qwen-Video-Edit repurposes an image editing DiT for video

AntonObukhov1 · x · 2026-09-10

Two related research drops: Marigold V2 (SIGGRAPH Asia 2026) upgrades the single-GPU depth-estimation post-training trick to a diffusion transformer with sharper edges and more versatility, while Qwen-Video-Edit enables instruction-based video editing without any video-pretrained model — it teaches Qwen-Image-Edit's DiT to edit Wan 2.1 video-VAE latents directly via small in/out projections, tiling latent frames into one virtual image with grid RoPE.

Trained with LoRA or full fine-tuning on Ditto-1M (source, edited, instruction) triplets and refined with Wan 2.2 denoising, it supports long-video multi-instruction editing and portrait videos natively.

Related event: Marigold V2 Released: Single-GPU DiT Depth Estimator Heading to SIGGRAPH Asia 2026(9 posts)→

Original post →

More from Multimodal

Multimodal channel →