PAMI anchors object motion to body parts for text-to-HOI, +14.5% contact recall
Chuqiao Li · hf · 2026-10-09
Paper PAMI tackles text-conditioned full-body human-object interaction (HOI) generation. Prior methods model human and object as separate trajectories and learn coupling implicitly, causing object drift, missed contact, and penetration.
Inspired by the Hough Transform, PAMI lets multiple body-part anchors vote for object motion: object motion is expressed relative to part anchors, with a PamiVAE learning an interaction latent space and decoding frame-wise weights to aggregate votes.
Generation is coarse-to-fine:
- PamiGen: coarse text-to-interaction in the structured latent space
- PamiRefiner: recursively resolves fine contact geometry via hybrid surface sensing (long-range probes + short-range sensors)
On InterAct, PAMI achieves 14.5% higher contact recall than prior SOTA with more faithful object-relative motion; ablations validate both the voting representation and refinement stage.
More from Multimodal
- Odyssey launches Odyssey-3, claims SOTA world model on Physics-IQ benchmark — Scobleizer · 2026-10-09
- Nano Banana 2.1 becomes Google's best image editing model across all 7 edit actions — ArtificialAnlys · 2026-10-09
- Nano Banana 2.1 Improves on All 9 Measured Image Capabilities, Biggest Gains in Layout and Anatomy — ArtificialAnlys · 2026-10-09
- Nano Banana 2.1 Pushes the Quality-Price Frontier: Three Ranks Higher at Half the Price — ArtificialAnlys · 2026-10-09
- Google's Nano Banana 2.1 Ranks #4 on Image Generation and Editing at Half Its Predecessor's Price — ArtificialAnlys · 2026-10-09
- Stanford's Level-of-Token Diffusion cuts image and video generation cost with multiresolution tokens — GordonWetzstein · 2026-10-09