Robot hand turns book pages: UVTA visual-tactile-action model hits 70% vs 29% baseline
CyberRobooo · x · 2026-10-10
BIGAI, Sharpa and Shanghai Jiao Tong University unveiled UVTA, a unified visual-tactile-action model that lets a robot hand reliably turn single book pages.
- Why it's hard: contact tasks on paper, fabric, switches and soft packs are where humanoids choke — objects slip, deform and block the camera view.
- Data collection: instead of farming touch on robots, humans wore a passive 22-DoF exoskeleton capturing hand motion, fingertip pressure, wrist pose and a wrist camera — 1,000 human demos per task plus 150 robot demos.
- Key idea: different hands, same contact physics. Human tactile demos and a smaller robot set are fused into a shared visual-tactile-action representation that also predicts the next touch.
- Results: 70% average success across five contact tasks, versus 29% for the best visual-tactile baseline they ran.
Related event: SJTU's UVTA Unifies Vision and Touch for Robotic Dexterity(3 posts)→
More from Embodied
- Xpeng shows Physical AI at Paris Motor Show, vision-only FSD heading to Europe — LinusEkenstam · 2026-10-10
- FIND preprint: robot picks its own weaknesses, lifts 8-task success from 55% to 71.9% — GeorgiaChal · 2026-10-10
- Musk: Digital Optimus beats Diablo halfway through with no APIs, just screen pixels — XFreeze · 2026-10-10
- Tongji team builds spiderweb-inspired 6-axis robotic 3D printer with self-supporting structures — lukas_m_ziegler · 2026-10-10
- Sony surgical robot's autonomous suturing demo sparks debate over medical AI liability — TansuYegen · 2026-10-10
- VioLA trains humanoid policies on 140M mostly-human frames, hitting 88.6% zero-shot manipulation — Michael_J_Black · 2026-10-10