GraphVid controls video generation with interaction graphs and cuts FID by 39.9%
PLAN-Lab · hf · 2026-07-24
What it is
GraphVid is a graph-conditioned image-to-video model that lets users control multi-object interactions through structured interaction graphs instead of brittle text prompts or manual motion tracks.
Why it matters
- Trajectory-based control does not scale well when scenes get crowded or objects overlap.
- GraphVid uses a semantic interface built around interaction graphs, making control more flexible and more precise for multi-subject scenes.
- The authors also release GraphVid-Bench, a large-scale interaction-centric dataset with structured relational annotations.
Results
- Despite using less training data and fewer trainable parameters than prior motion-control methods, GraphVid reports strong controllability and video quality.
- Compared with Motion-I2V, it claims up to 39.9% lower FID and 37.6% lower FVD, plus better PSNR (9.87 → 15.98) and SSIM (0.38 → 0.61).
- The paper argues that structured semantic interfaces are a promising direction for controllable video generation.
More from Multimodal
- New ComfyUI node automates image upscaling with context-anchored tiles — blakeem · 2026-07-24
- Black Forest Labs video models are being used to automate Audi manufacturing — hsu_byron · 2026-07-24
- Unlimited-OCR returns to No. 1 on Hugging Face trending models — _akhaliq · 2026-07-24
- LTX 2.3 camera motion in ComfyUI is causing severe blur and ghosting — Alarming_Watch9109 · 2026-07-24
- ChatGPT Images still wins on niche hairstyle ideas, despite weak human-photo results — flowersslop · 2026-07-24
- AI-Driven Promo Video Creation: Major Upgrades in Animation Camera Control — AlchainHust · 2026-07-24