COSMI composes single-object captures into 222k multi-object interaction sequences, 30x larger than prior sets
UniTuebingen · hf · 2026-10-07
University of Tübingen researchers released COSMI, addressing the scarcity of multi-object human-interaction data.
Key insight: interactions are local, so single-object captures already contain the building blocks of multi-object activities. The team composes contact-consistent clips, mirrors them for hand balance, transfers them across bodies, and filters pairings with a language model plus geometric checks — so the dataset grows combinatorially with clips rather than recording time.
Results:
- 222k sequences, 275 hours, up to five objects — nearly 30x the largest existing multi-object capture
- Trained a text-to-interaction diffusion transformer with weight-shared object slots handling a variable number of objects
- Outperforms baselines on text alignment and contact accuracy, with the largest margin on unseen objects
Code, models and the dataset pipeline will be open-sourced.
More from Embodied
- Newsmax attacks Comma.ai's 'DIY self-driving', evoking early Tesla FSD panic — walkingriver · 2026-10-07
- Health AI startups raised ~$768M this week, led by Devoted Health's $555M — HealthcareAIGuy · 2026-10-07
- 'Young inventor' AI glasses exposed as 189 yuan Alibaba white-label resell — found by Claude — lxfater · 2026-10-07
- Sesame announces AI eyewear line, Made in Japan and launching in 2027 — testingcatalog · 2026-10-07
- Tesla's Cybercab rests on aggressive FMVSS interpretation NHTSA could reject — binarybits · 2026-10-07
- Waymo Colors Inside FMVSS Lines; Tesla Bets on Regulatory Favor, Analyst Argues — binarybits · 2026-10-07