COSMI composes single-object HOI datasets into 222k multi-object sequences to train a diffusion transformer
CSProfKGD · x · 2026-10-08
Gerard Pons-Moll's group at Tübingen introduces COSMI, built on the observation that most human-object interactions are local (a hand holds a cup, a chair supports the pelvis), so compositionality beats brute-force capture.
- Uses LLM reasoning plus geometric checks to compose existing single-object datasets into 222k multi-object sequences totaling 275 hours, with up to 5 objects per sequence
- Trains a single diffusion transformer for interactions with 1-5 objects, predicting each object relative to its interacting body part, generalizing to unseen objects and combinations
- The related AvaImg project (ECCV 2026) produces high-fidelity SMPL-X+D registrations with UV texture and displacement maps, fixing penetration artifacts present in dataset ground truth; a Blender add-on unifying 15 HOI datasets is also released
Related event: COSMI: 222K Multi-Object Interaction Sequences via Compositional Synthesis(2 posts)→
More from Research
- ExploreNet Boosts Diffusion GRPO by 14% by Learning Where to Explore — StellaLisy · 2026-10-08
- ExploreNet ablation: perturbing learned high-sensitivity channels drives larger human-perceived change — StellaLisy · 2026-10-08
- ExploreNet's gain comes from targeted exploration, not larger noise, ablation shows — StellaLisy · 2026-10-08
- ExploreNet diverges from SD-3.5 more often (86.1%) and beats FlowGRPO on benchmarks — StellaLisy · 2026-10-08
- ExploreNet: learning targeted exploration noise for GRPO in flow-matching models — StellaLisy · 2026-10-08
- ExploreNet outperforms FlowGRPO by learning state-dependent exploration noise — StellaLisy · 2026-10-08