PKU's ViGAR Hierarchical World-Action Model Boosts Robot Manipulation Success by 12.86 Points
DAGroup-PKU · hf · 2026-10-08
Peking University's DA Group proposes ViGAR, a hierarchical world-action model for long-horizon compositional robot manipulation.
- Architecture: a visual subgoal planner predicts the next subtask's visual subgoal from the observation and instruction; a subgoal executor then jointly generates future visual trajectories and actions conditioned on that subgoal. Both share a pretrained world-model representation.
- In-context learning: a single goal image can induce different subtask decompositions and behaviors without parameter updates.
- Results: 82.00% and 67.02% success on RoboTwin Clean2Random (Clean/Random), beating the strongest baseline by 12.86 points on average; validated on five real-world compositional and two in-context tasks.
More from Embodied
- Developer uses Codex for CAD and build docs in physical robot hardware project — OpenAIDevs · 2026-10-08
- Twitter user dance-battles a Unitree robot in real life — zealcaiden · 2026-10-08
- Nvidia's DreamDojo world model wins ICML spotlight despite critical bugs in released code — Amazing-Fox-7295 · 2026-10-08
- Vacuum robot caught running away from home, poster says they get it — BotLordBotflix · 2026-10-08
- Bulkhead drone motor design concept targets 60k motors per month in production — MatthewChang · 2026-10-08
- This robot vacuum ran away from home and honestly, we get it — DankestMage99 · 2026-10-08