Being-H0.8 adds tactile prediction to robot world models with 500,000 hours of video
量子位 · wechat · 2026-07-28
BeingBeyond released Being-H0.8, an implicit tactile world-action model built from human video data.
Key points:
- The model adds tactile signals to the latent world model so robots can predict contact, not just motion.
- It uses more than 500,000 hours of first-person video data, plus a large data pipeline for filtering, motion recovery, and language annotation.
- TactoHand reconstructs contact and proximity labels from video, while pressure-sensor data adds physical force information.
- A universal tactile encoder normalizes different tactile modalities into shared tokens.
- TopoHand aligns human hands, dexterous hands, and grippers into a unified action representation.
In real robot tests, the system handled bag retrieval, calligraphy, toothpaste squeezing, and snack grasping — all tasks that depend heavily on touch feedback.
More from Embodied
- Dex-Net in 60 Seconds: A Classic Look at Robotic Grasping Uncertainty — berkeley_ai · 2026-07-28
- Webots, an open-source robot simulator for vehicles and mechanical systems — tom_doerr · 2026-07-28
- Tesla's New Feature: Headlights Project Navigation Path Directly onto the Road — sven_ai · 2026-07-28
- Robots Don't Need to Beat Humans, Just Fill the Empty Shift — VraserX · 2026-07-28
- Zhigu Tianchu raises nearly RMB 100 million as commercial cooking robots enter the monetization phase — 创业邦 · 2026-07-28
- Real-Time Speech-to-Action Mapping Processed Entirely On-Robot — tctjr · 2026-07-28