ByteDance’s FlowMimic trains video editing without masks by mimicking image and video modalities

ByteDance-Seed · hf · 2026-07-21

ByteDance’s FlowMimic aims to unify video editing and generation without masks

FlowMimic explores a single model that can handle both generation and editing for image and video modalities. The main bottleneck it targets is data collection: existing video editing pipelines often require manual mask annotation, synthetic pair generation, and heavy VLM-based filtering, which makes the task space narrow and hard to scale.

Core ideas

Training design

Original post →

More from Multimodal

Multimodal channel →