Qwen-Video-Edit: zero-training projection layers let image editors edit video latents

luisdans · x · 2026-09-04

The Qwen team (Byp215Bai) open-sourced Qwen-Video-Edit with a technical blog. Core finding: video and image latent spaces are far less different than almost everyone assumes — an image editing model that has never seen a video can directly edit the latents of a video generation model through two zero-training projection layers.

The difference isn't zero, though — it hides in one specific place: temporal compression. The post walks through the full chain of experiments used to measure it.

The project is on GitHub, integrated into DiffSynth-Studio, and ships with ComfyUI nodes and a full install guide.

Related event: Alibaba Open-Sources Qwen-Video-Edit: Zero-Training Bridge Between Image and Video Latent Spaces(2 posts)→

Original post →

More from Multimodal

Multimodal channel →