Tsinghua's EMERGE-Policy Orchestrates VLA, World Models as Multi-Agent Tools, Hits 94% Under Real-World Disturbances

机器之心 · wechat · 2026-09-12

Tsinghua SIGS and collaborators released EMERGE-Policy, an open-source multi-agent framework for robot manipulation (paper on arXiv). Instead of a single end-to-end VLA, it wraps VLA (pi0.5), motion primitives, world models (CosmosPolicy) and verifiers into a unified skill library orchestrated by graph-structured agents, letting system-level intelligence emerge from component collaboration.

Architecture

Results: 99.2% on standard LIBERO (+0.7% over base); 93.9% under LIBERO-Plus perturbations (+11.7% over the CosmosPolicy base); strong RoboDojo-Sim memory scores. On a real-robot multi-layer cup-stacking task, it sustains 85%-94% success under human sabotage and cup/visual changes.

Original post →

More from Embodied

Embodied channel →