PAMI anchors object motion to body parts for text-to-HOI, +14.5% contact recall

Chuqiao Li · hf · 2026-10-09

Paper PAMI tackles text-conditioned full-body human-object interaction (HOI) generation. Prior methods model human and object as separate trajectories and learn coupling implicitly, causing object drift, missed contact, and penetration.

Inspired by the Hough Transform, PAMI lets multiple body-part anchors vote for object motion: object motion is expressed relative to part anchors, with a PamiVAE learning an interaction latent space and decoding frame-wise weights to aggregate votes.

Generation is coarse-to-fine:

On InterAct, PAMI achieves 14.5% higher contact recall than prior SOTA with more faithful object-relative motion; ablations validate both the voting representation and refinement stage.

Original post →

More from Multimodal

Multimodal channel →