RoboCoach uses world-model-imagined failures as coaching: 13.3% to 75% success with 150 demos

Tsinghua · hf · 2026-10-01

Tsinghua's RoboCoach treats world models as active coaches for long-horizon robot manipulation. Its Route-Imagine-Diagnose-Improve loop runs reusable skill experts inside a shared action-conditioned world model, using a progress judge to log the first failing subtask; aggregated records decide which demonstrations to collect and which expert adapters to update.

Imagined and deployed success correlate at rho = 0.840 across 22 task-policy pairs on two sim suites and two real robots. With just 150 extra subtask demonstrations, success rises from 13.3% to 75.0% on Franka and 40.0% to 83.8% on AgileX; coached experts transfer to four held-out compositions at 35.0% average success versus 0% for a uniformly-updated shared-policy baseline.

Original post →

More from Embodied

Embodied channel →