ThinkV2V makes the MLLM think before editing: a 5B DiT beats 14B baselines on implicit instructions

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Donghao Zhou, Haoyang He, Fan Zhang, Hao Yang, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, Jie Wang, Qiangpeng Yang, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng

cs.CV

2026-09-30

ThinkV2V turns the MLLM into an explicit thinker whose chain-of-thought drives a 5B DiT; the 5B model outscores 14B baselines on both editing benchmarks.

What problem this solves

Real editing instructions rarely name their target directly. "Remove the container used for carrying the food for this outdoor meal" requires the model to infer that the container is a picnic basket before anything can be edited. Current instruction-guided video editors (OmniVideo, ICVE, OpenVE-Edit, and peers) use the MLLM as a semantic encoder: it jointly embeds the instruction and the source video, hands features to the generator, and stops at perception-level alignment. When the real intent sits in causal or contextual cues, the model latches onto surface keywords and edits the wrong object or region. The paper's diagnosis is blunt: the bottleneck is the missing think-before-edit step, not generation quality. Datasets and benchmarks share the same debt, testing almost exclusively direct instructions.

Method

Three architectural parts plus a training and inference recipe.

Training runs a three-stage curriculum: 480p basic edits on OpenVE-HQ-1M (2 epochs), 720p adaptation (0.5 epoch), then 720p reasoning-intensive tuning on ThinkV2V-150K (1.5 epochs); only the connector and DiT are trained. High resolution and complex instructions introduced together make spatial-temporal alignment and reasoning-driven execution fight each other, so they are separated. At inference, refined prompts are fed back for multiple rounds of serial refinement, then the MLLM picks the best of 8 candidates against the original intent. Iteration without selection compounds early misreadings.

Data comes from OpenVE-3M: five edit categories retained, filtered by Gemini-2.5-Flash scoring plus CLIP-F and TF temporal metrics into a 1M high-quality pool, from which 150K pairs get their direct instructions rewritten into reasoning-oriented ones by Gemini-2.5-Pro. ThinkV2V-Bench rewrites OpenVE-Bench the same way, yielding 308 pairs.

Results

In the main table, ThinkV2V ranks first on overall across ThinkV2V-Bench and OpenVE-Bench under both judges (Seed-1.6-VL and Gemini-2.5-Pro), with a 5B DiT beating VACE (14B), DITTO (14B), and ICVE (13B). Not a clean sweep: on ThinkV2V-Bench, DITTO leads Global Style 3.28 to 3.13, and ReCo is strongest on Local Remove. Gains carry over to the direct-instruction OpenVE-Bench, with Local Change and Local Remove the standout categories.

Ablations (Seed-1.6-VL judge, overall score):

ConfigOverall
MLLM as static encoder, no thinking2.51
Thinking, all-token features2.63
Thinking, answer-token features2.68

Answer tokens alone beat all generated tokens: the reasoning trace adds redundancy, not signal. Curriculum order:

Training data orderOverall
Complex only2.10
Complex-to-simple2.41
Mixed2.47
Simple only2.52
Simple-to-complex (ThinkV2V)2.68

On inference-time scaling, serial refinement alone lands at 2.67, which the paper states brings no gain; adding best-of-8 selection brings 2.72. Iteration without selection is not enough.

Why it matters

First, the engineering details of wiring a reasoning model into a diffusion model are fully specified: which hidden states to tap, how to design the connector, how to combine conditions. Anyone building MLLM-to-DiT pipelines can lift this directly. Second, 5B beating 14B backs the core claim that the editing bottleneck is semantic understanding, not generation capacity. Third, best-of-N pays off without retraining, because the MLLM's explicit text output can be reused and screened, a dividend unavailable when the MLLM is just an encoder. The architecture remains the familiar MLLM-plus-DiT combination; the real increment is making thinking an explicit, trainable, selectable signal. A solid systems paper, not a new paradigm.

Limitations

The authors list three in Appendix C: explicit thinking raises training and inference cost and latency; thinking content is never independently evaluated for correctness, only downstream editing outcomes; and dataset plus benchmark coverage of longer-horizon, open-domain, rarer edits is thin.

More question marks on a close read. The "implicit instructions" in both training data and benchmark are all synthesized by Gemini models, and whether that distribution matches how real users phrase requests is untested. Evaluation rests entirely on two MLLM judges with no human study, and a judge's taste for verbose rewrites may leak into scores. ThinkV2V-Bench has only 308 pairs. The "5B" counts the DiT alone; an 8B MLLM sits in front, and inference adds serial refinement plus best-of-8, so the total bill is far from small.

Terms

Source

Related papers

All paper explainers