MPIE-Bench: Evaluating Anatomical Errors in Multi-Person Image Editing
muset-ai · hf · 2026-07-31
While text-to-image models excel at single-subject synthesis, generating multi-person physical interactions (like embraces or carries) often results in fused limbs and interpenetrating bodies. Existing evaluations and VLM judges frequently overlook these anatomical flaws.
To address this, researchers introduced MPIE-Bench:
- Dataset: 2,500 video-mined editing triplets across 405 scenes, 14 interaction categories, and four contact densities.
- New Axes: Proposes MPIE-Eval, leveraging multi-person mesh reconstruction to score contact-time geometry across Anatomy (completeness of bodies) and Interaction (penetration/surface distance matching).
- Findings: Across ten editors, no single model excelled on both axes (topping at 0.65 and 0.72), while VLM judges saturated above 0.95. The new axes correlate much closer to human judgment.
More from Multimodal
- Hailuo AI Demonstrates 15-Second Video Generation from a Single Image — LudovicCreator · 2026-07-31
- MiniMax H3 Video Model Enters Chatbot Arena, Open Weights Coming Soon — arena · 2026-07-31
- Mind-Blowing AI Concept: Visualizing Pigeon 'Flight-Mile' Delivery to Dodge Traffic — sivakumaranvlsi · 2026-07-31
- MiniMax H3 Video Model Now Available on Vercel AI Gateway — evilrabbit_ · 2026-07-31
- Testing Flux 3: Generating Synchronized Split-Screen Videos via Complex Prompts — umesh_ai · 2026-07-31
- RefCaptioner: Grounding Video Captions to Multiple Reference Images — KlingTeam · 2026-07-31