Peking University Proposes New HOI Image Editing Framework
jiqizhixin · x · 2026-07-19
Researchers from Peking University have proposed a new image editing method specifically designed to solve the challenges of editing complex Human-Object Interaction (HOI).
- Core Idea: Repurposing a Image-to-Video (I2V) generation model to generate dynamic interaction sequences, and introducing a self-correction loop (SCPE) to iteratively refine the prompt.
- Performance: In the complex HOI-Edit benchmark, this method's editing capabilities rival those of top-tier image editors like Nano Banana.
More from Multimodal
- HOMIE pairs Qwen3-VL-2B with Wan2.1 for human-object-centric video personalization — switch2stock · 2026-07-21
- Video Models Cut Ad Production Costs by 90-99%: Runway Enterprise Data — c_valenzuelab · 2026-07-21
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21