CoinVE-Edit: apply up to 5 video edit instructions in a single forward pass

pmttyji · reddit · 2026-08-20

CoinVE-Edit, newly released on Hugging Face, is a compositional video-editing model built on Wan2.1-T2V-14B (video DiT) and Qwen3-VL-8B (MLLM encoder). It handles multi-instruction editing: 2–5 edit instructions in a single forward pass, each confined to its designated region via per-instruction mask injection through a lightweight mask head.

It supports Replace, Add, Remove, and Background Change in any combination while keeping overall video coherent. Trained on the CoinVE-200K dataset of 200K+ rigorously filtered video-edit pairs.

Related event: Tencent Releases CoinVE-Edit Model and CoinVE-200K Dataset(2 posts)→

Original post →

More from Multimodal

Multimodal channel →