CoinVE-Edit: apply up to 5 video edit instructions in a single forward pass
pmttyji · reddit · 2026-08-20
CoinVE-Edit, newly released on Hugging Face, is a compositional video-editing model built on Wan2.1-T2V-14B (video DiT) and Qwen3-VL-8B (MLLM encoder). It handles multi-instruction editing: 2–5 edit instructions in a single forward pass, each confined to its designated region via per-instruction mask injection through a lightweight mask head.
It supports Replace, Add, Remove, and Background Change in any combination while keeping overall video coherent. Trained on the CoinVE-200K dataset of 200K+ rigorously filtered video-edit pairs.
Related event: Tencent Releases CoinVE-Edit Model and CoinVE-200K Dataset(2 posts)→
More from Multimodal
- Dreamina Fun Demo: Creating 'Around the World' Visual Effects with Prompt — umesh_ai · 2026-08-20
- H3TiledLoopSpaceTime Demonstrates H3-Based Spatiotemporal Video — DuHal9000 · 2026-08-20
- MiniMax H3 lands on Magnific: mix 9 images, 3 videos and 3 audio refs in one prompt — AIwithGhotai · 2026-08-20
- MiniMax H3 excels at handling long sequences in single prompts for video generation — egeberkina · 2026-08-20
- each::labs releases Video API with 39 models for post-processing and packaging — thetripathi58 · 2026-08-20
- AI Video Tools Review: Runway leads realism, Luma excels in motion — Acrobatic_Show_9092 · 2026-08-20