World Synesthesia Model Enables 1-Minute Noise-Robust In-Hand Rotation, CoRL 2026 Paper
机器之心 · wechat · 2026-09-17
Sharpa Robotics researchers propose the World Synesthesia Model (WSM), a CoRL 2026 (high-score) work tackling why dexterous in-hand manipulation policies collapse under real-world depth noise, sparse touch, and perturbations.
- Core idea: unify visual geometry, tactile contact, proprioception, and action history into a Dreamer-style RSSM world-model state, letting the policy continuously infer the latent physical state of the hand-object system. Training injects noisy depth as input but supervises clean depth, forcing the recurrent state to recover usable geometry.
- Ablations: removing recurrent input drops return to 705.4; frame-wise latents to 667.6; full WSM reaches 753.3 — clean-geometry supervision plus action-conditioned recurrence is what drives the gain.
- Reusable prior: a WSM pre-trained on 9 objects transfers to 49 novel objects, lifting average rotation from 3.28 to 9.37 rad/episode and cutting drop rate from 6% to 0.3%.
- Real deployment: on a 22-DoF five-finger hand, the policy achieves multi-object z/y-axis rotation, zero-shot generalization, 3-6s perturbation recovery, and 1+ minute stable rotation of a real cube, beating open-loop and tactile-only baselines.
The key takeaway: world models need not "dream the future" — learning a better representation of the present suffices to close the sim-to-real gap for dexterous manipulation.
More from Embodied
- Figure Teases New Robotics Announcement, 'Robotics Will Accelerate Today' — ChrisGPT · 2026-09-17
- Waymo says it simulates elephants on freeways, gets memed for real-world failures — RexDouglass · 2026-09-17
- Sofía Dudas Kicks Off SSAD 2026 Day 3 With Talk on World Models for Autonomous Driving — abursuc · 2026-09-17
- Ropedia launches academic partner program offering HOMIE Gen2 kits for physical AI data — MengdiWang10 · 2026-09-17
- Doubao launches in-car cockpit assistant with multimodal vehicle control — xiaohu · 2026-09-17
- ActionPiece rethinks action tokenization for VLA models, hits 94.8% on LIBERO — DeepCybo · 2026-09-17