GraphVid controls video generation with interaction graphs and cuts FID by 39.9%
PLAN-Lab · hf · 2026-07-24
What it is
GraphVid is a graph-conditioned image-to-video model that lets users control multi-object interactions through structured interaction graphs instead of brittle text prompts or manual motion tracks.
Why it matters
- Trajectory-based control does not scale well when scenes get crowded or objects overlap.
- GraphVid uses a semantic interface built around interaction graphs, making control more flexible and more precise for multi-subject scenes.
- The authors also release GraphVid-Bench, a large-scale interaction-centric dataset with structured relational annotations.
Results
- Despite using less training data and fewer trainable parameters than prior motion-control methods, GraphVid reports strong controllability and video quality.
- Compared with Motion-I2V, it claims up to 39.9% lower FID and 37.6% lower FVD, plus better PSNR (9.87 → 15.98) and SSIM (0.38 → 0.61).
- The paper argues that structured semantic interfaces are a promising direction for controllable video generation.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11