Tsinghua's EMERGE-Policy Orchestrates VLA, World Models as Multi-Agent Tools, Hits 94% Under Real-World Disturbances
机器之心 · wechat · 2026-09-12
Tsinghua SIGS and collaborators released EMERGE-Policy, an open-source multi-agent framework for robot manipulation (paper on arXiv). Instead of a single end-to-end VLA, it wraps VLA (pi0.5), motion primitives, world models (CosmosPolicy) and verifiers into a unified skill library orchestrated by graph-structured agents, letting system-level intelligence emerge from component collaboration.
Architecture
- A Main Agent decomposes long-horizon tasks and schedules via hierarchical context engineering; Perception/Verification/Monitor sub-agents run in isolated contexts.
- Three file-level memories: PLAN.md (task graph), HISTORY.md (observations), MEMORY.md (environment facts), auto-compressed when context overflows.
- The world model serves as a simulator to rehearse candidate action paths before execution, scored by verifiers — "think before acting."
Results: 99.2% on standard LIBERO (+0.7% over base); 93.9% under LIBERO-Plus perturbations (+11.7% over the CosmosPolicy base); strong RoboDojo-Sim memory scores. On a real-robot multi-layer cup-stacking task, it sustains 85%-94% success under human sabotage and cup/visual changes.
More from Embodied
- FoloToy's AI Passport toy sells out as buyers wait for shipment — ezshine · 2026-09-12
- RLWRLD Lands 5 CoRL 2026 Papers; Human-Video Pretraining Lifts VLA Success to 80.3% — chris_j_paxton · 2026-09-12
- Robotics Researcher Satirizes How to Fake a Robotics Paper Ahead of ICRA Deadline — chris_j_paxton · 2026-09-12
- Trust3R: ICML 2026 Paper Adds Evidential Uncertainty to Feed-Forward 3D Reconstruction — rsasaki0109 · 2026-09-12
- BlackBerry QNX: the overlooked 'Physical AI tollbooth' positioned for a decade-defining comeback — pdamodaran · 2026-09-12
- A Simulated Fruit Fly Tries to Solve a Rubik's Cube — Matthew Berman · 2026-09-12