COSMI composes single-object captures into 222k multi-object interaction sequences, 30x larger than prior sets

UniTuebingen · hf · 2026-10-07

University of Tübingen researchers released COSMI, addressing the scarcity of multi-object human-interaction data.

Key insight: interactions are local, so single-object captures already contain the building blocks of multi-object activities. The team composes contact-consistent clips, mirrors them for hand balance, transfers them across bodies, and filters pairings with a language model plus geometric checks — so the dataset grows combinatorially with clips rather than recording time.

Results:

Code, models and the dataset pipeline will be open-sourced.

Original post →

More from Embodied

Embodied channel →