Qwen-Video-Edit enables instruction-based video editing by repurposing an image model
AgeNo5351 · reddit · 2026-08-24
Core Mechanism
Qwen-Video-Edit enables instruction-based video editing by adapting Qwen-Image-Edit's Transformer to operate directly on video-VAE latents.
Technical Details
- Cross-Space Mapping: Uses two tiny projection layers to bridge Wan 2.1's video latent space into the DiT's token space, allowing static video frames to be embedded exactly like images.
- Tile Arrangement: Latent frames are arranged as tiles of one large virtual image, maintaining the positional encoding treatment from the image model's pretraining.
- Training Pipeline: Fine-tuned with LoRA or full parameters on the Ditto-1M dataset (source, edited, instruction triplets), then refined with a few steps of Wan 2.2 denoising-enhancement.
Availability
Model weights, code, and a demo page are available.
More from Multimodal
- Grok demonstrates video editing with automatic music sync — XFreeze · 2026-08-24
- Circuit Board City generative art demo and open source code — jasonkneen · 2026-08-24
- Trimming Video Latent in Stage 2 Causes Flickering — parth0202 · 2026-08-24
- Fix Low-Res MiniMax H3 Generations with Ultimate SD Upscale — alisitskii · 2026-08-24
- Kimi K3 Generates Jellyfish-Inspired Deep Sea Robot Design — mishig25 · 2026-08-24
- Minimax Character Swap Workflow: The Green Dummy Strategy — lhg31 · 2026-08-24