UniWAM unifies world model and action policy, hits SOTA on LIBERO and RoboTwin benchmarks
LexiLove · x · 2026-10-05
UniWAM (HKUST-GZ, OLA Dimensions, CMU, PKU and others) stops separating 'understand, predict, act': one architecture combines a VLM-based physical reasoner, a video world model, and an action policy.
- Trained on robot demos, human egocentric video, and VQA data with a rigorous cleaning/annotation pipeline
- SOTA results: 99.2% LIBERO, 92.6% LIBERO-Plus, 68.3% RoboTwin Clean2Rand; real-world demos include pick-anything, bimanual drawer storage, and long-horizon tidying
- Key finding: human+robot co-training follows a log-linear scaling trend, suggesting human video matters a lot for scaling robot foundation models
Paper, code, and checkpoints are public.
More from Embodied
- Panasonic building its own humanoid robot, targeting 2029 production in its own factories — CyberRobooo · 2026-10-05
- FailBank Turns Runtime Shield Feedback into Persistent VLA Policy Gains (+25.4 Points Success) — notredame · 2026-10-05
- Reka releases Rho-1 research preview: a 19B omni model unifying text, images, video and robot actions — RekaAILabs · 2026-10-05
- NVIDIA's ENPIRE lets coding agents self-improve robot policies in the real world, hitting 99% success — chris_j_paxton · 2026-10-05
- Sensori demos humanoid Yuri: natural voice chat plus VR and exoskeleton teleoperation — YuXiang_IRVL · 2026-10-05
- EyeRobot 2.0: active gaze lifts robot success from 52% to 67% — no wrist cameras needed — stepjamUK · 2026-10-05