Marigold V2 arrives as Qwen-Video-Edit repurposes an image editing DiT for video
AntonObukhov1 · x · 2026-09-10
Two related research drops: Marigold V2 (SIGGRAPH Asia 2026) upgrades the single-GPU depth-estimation post-training trick to a diffusion transformer with sharper edges and more versatility, while Qwen-Video-Edit enables instruction-based video editing without any video-pretrained model — it teaches Qwen-Image-Edit's DiT to edit Wan 2.1 video-VAE latents directly via small in/out projections, tiling latent frames into one virtual image with grid RoPE.
Trained with LoRA or full fine-tuning on Ditto-1M (source, edited, instruction) triplets and refined with Wan 2.2 denoising, it supports long-video multi-instruction editing and portrait videos natively.
More from Multimodal
- GPT-6 Astra drives Houdini for hands-off procedural modeling with one prompt — ssh4net · 2026-09-10
- Gradium, from Kyutai's lab, launches Voice Design: prompt-to-voice in seconds, free in API — RemiCadene · 2026-09-10
- Full prompt shared: recreating early-2000s DV camcorder realism in a 30s Seedance 2.5 video — SimplyAnnisa · 2026-09-10
- New Midjourney V8.2 style showcased in a morning-generation share — azed_ai · 2026-09-10
- Midjourney Meets Hailuo H3: A Proven Image-to-Video Workflow for MJ-Style Motion — techhalla · 2026-09-10
- Flova's AI Video Turns an iPhone Unboxing Into an Apple-Commercial Lookalike — iamfakhrealam · 2026-09-10