Qwen-Video-Edit repurposes an image editing DiT to edit videos, no video-pretrained backbone needed

qixing_huang · x · 2026-08-18

Qwen-Video-Edit introduces instruction-based video editing by repurposing the pretrained image editing model Qwen-Image-Edit's DiT to operate directly on Wan 2.1 video-VAE latents — no video-pretrained transformer required.

Code and the Hugging Face model are released.

Original post →

More from Multimodal

Multimodal channel →