ThinkV2V: Reasoning-Driven Video Editing — a 5B Model Beats 10B Baselines on Complex Instructions
Donghao Zhou · hf · 2026-10-01
- Existing instruction-guided video editors use MLLMs mainly as semantic encoders, struggling with implicit edits requiring causal/semantic reasoning. ThinkV2V activates explicit MLLM thinking before visual generation.
- Built on an MLLM-to-DiT architecture that turns reasoning over the source video and instruction into refined conditioning signals, it combines Progressive Curriculum Training (basic editing → reasoning-intensive cases) with Inference-Time Thinking Scaling (iteratively refining candidate prompts and selecting the most reliable one).
- The team releases the ThinkV2V-150K dataset and ThinkV2V-Bench; their 5B DiT substantially outperforms larger 10B baselines, achieving SOTA on both complex and standard editing scenarios.
More from Multimodal
- Seedance 2.5 demo brings anime-level dual-sword choreography into photorealistic cinema — SimplyAnnisa · 2026-10-01
- Midjourney --sref 3896456162 recreates 1970s Kodak film & disco aesthetics — michaelrabone · 2026-10-01
- A finished LoRA run doesn't mean it learned the style: a reproducible SDXL validation workflow — no3us · 2026-10-01
- Editor open-sources open-fusion-mcp: Claude builds editable motion graphics inside DaVinci Resolve — JohnnyLegion · 2026-10-01
- AI-generated clip: Raven Noire writing in her diary while listening to goth music — Street-Pound5762 · 2026-10-01
- Failed MiniMax H3 VAE detail experiment yields a useful 2X detail VAE and workflow — NoMouse9610 · 2026-10-01