ALIVE Makes Inserted Video Objects Interact, Beating Baselines by 43.9%
university-of-rochester · hf · 2026-10-07
University of Rochester presents ALIVE, a first-frame-guided video editing framework that makes inserted objects participate in interactions like being picked up or manipulated.
Method
- Needs only an edited first frame plus an instruction naming the added object
- Curates 35,800 editing pairs mixing 3D-rendered, model-generated, and real videos with ROSE general editing pairs
- Trains a VLM to predict interaction guidance from the same inputs, no extra user input required
- Introduces the ALIVE-interaction benchmark using a unified VLM-based protocol
Results: Without VLM guidance, ALIVE improves Overall over the strongest baseline by 43.9% and 4.4% on two benchmarks; VLM-predicted guidance adds another 0.95 points.
More from Multimodal
- Tencent Hunyuan releases WorldPlay2, a consistent interactive world model steerable by prompts — Scobleizer · 2026-10-07
- 60 Minutes aired an AI version of correspondent Jon Wertheim, then promised no more AI content — c_valenzuelab · 2026-10-07
- Spider-Man vs X-Men Doomsday battle imagined with Seedance 2 — theNerdSoul · 2026-10-07
- Dreamina Canvas adds Extract Motion: copy the camera move, not the shot — Div_pradeep · 2026-10-07
- Blender's City Generator 2.0 builds full 3D metropolises in seconds — Scobleizer · 2026-10-07
- Runway unveils 'What We Forgot,' an AI film set in 2030 looking back at pre-AI life — c_valenzuelab · 2026-10-07