UVTA paper, code and dataset open-sourced: shared visual-tactile-action representation from human demos
CyberRobooo · x · 2026-10-10
The Shanghai Jiao Tong AI school team published the UVTA project page with paper, code and dataset all open.
- Motivation: dexterous manipulation needs touch, but collecting tactile demos on robot hands doesn't scale; human interaction is the scalable source of contact experience.
- System: a passive 22-DoF exoskeleton records wrist RGB, 20-D touch and 31-D action at 30Hz; 1,000 human demos per task plus 150 robot demos, dataset on Hugging Face.
- Model: a 384-D visual token and 20-D tactile token fuse into a 404-D interaction token, with a diffusion action head emitting 16-step action chunks and future-tactile prediction as auxiliary training.
- The pitch is cross-embodiment transfer: hands change, contact physics doesn't.
Related event: SJTU's UVTA Unifies Vision and Touch for Robotic Dexterity(3 posts)→
More from Embodied
- UniAD veterans launch Archon Robotics to build whole-body humanoid foundation model in China — micoolcho · 2026-10-10
- "She's a 10/10, but she's a robot": humanoid demo turns heads — creatoroff · 2026-10-10
- DepthART: A 6M-Parameter Depth Model Keeps Foundation-Model Generalization on Edge Devices — jiqizhixin · 2026-10-10
- ModelBest partners with Lumming Robotics to deploy on-device LLMs for embodied AI — 面壁智能 · 2026-10-10
- Xpeng shows Physical AI at Paris Motor Show, vision-only FSD heading to Europe — LinusEkenstam · 2026-10-10
- FIND preprint: robot picks its own weaknesses, lifts 8-task success from 55% to 71.9% — GeorgiaChal · 2026-10-10