World Action Agent Hits 75.6% on LIBERO-Pro via VLM Action Rehearsal
Yehang Zhang · hf · 2026-09-25
World Action Agent (WAA) lets VLMs pilot robots through a multi-agent harness with a visual action workspace.
- Design: Contact views auto-selected from scene geometry; action rehearsal turns each action into an editable proposal previewed by the agent or an Imagination Agent; in-view correction closes the loop between observation, rehearsal, and execution
- Skill acquisition: Multimodal skills evolve from expert videos and human teaching under evidence-based review; interaction traces fine-tune smaller VLMs to pilot the same harness
- Results: With skills evolved only from LIBERO-90, WAA reaches a SOTA 75.6% average success on LIBERO-Pro, outperforming end-to-end VLAs, code-as-policy, and a same-backbone visual harness; skills transfer to robosuite, and fine-tuning Qwen3.5-9B on traces lifts out-of-domain success from 1.7% to 43.3%
More from Embodied
- Waymo: 82% Fewer Injury Crashes Than Humans Across 271 Million Driverless Miles — reed · 2026-09-25
- HRI 2027 Adds Archival Industry White Paper Track for Real-World Robot Deployments — petitegeek · 2026-09-25
- Google dev imagines orchestrating AI agents via smart glasses: tap their shoulder, watch their screen — jason_mayes · 2026-09-25
- Climbing analysis in 3D with iPhone LiDAR: open-source SAM 3.1 + ViTPose demos — dosco · 2026-09-25
- Menlo open-sources Asimov 1 humanoid locomotion policy and Isaac Lab training code — freelerobot · 2026-09-25
- Human-Robot Dialogue workshop at IROS 2026: NVIDIA, MIT, Georgia Tech speakers lined up — dhadfieldmenell · 2026-09-25