Star Era's VPP2 Tops RoboDojo, Beating GPT-6-Astra by Nearly 10 Points

量子位 · wechat · 2026-10-09

Chinese embodied-AI company Star Era (backed by Tsinghua) has topped the RoboDojo simulation leaderboard with its World Action Model VPP2, posting 32.26% average success rate and 39.26 average score — both first place, ahead of GPT-6-Astra (22.48%/28.97) by 9.78 points and 10.29 points, and beating Physical Intelligence π0.5 and Nvidia GR00T-N1.7. RoboDojo, led by HKU MMLab with nearly 20 institutions, spans 42 dual-arm manipulation tasks across generalization, precision, long-horizon, memory and open-vocabulary instruction. VPP2 also ranks first in generalization, precision and memory, without extra data or augmentation.

Technically, VPP2 decouples video prediction from action learning in three stages: event-level video pretraining, fixed 8-second clip post-training with consistency distillation (0.12s to predict 8s of visual change), and a 0.9B ActionDiT expert trained with the video model frozen and LoRA-adapted, giving 0.22s total action-chunk latency. On real ALOHA robots in zero-shot tests, VPP2 averaged 58.5% success across 10 tasks vs π0.5's 40%, best in 9 of 10; LIBERO-Pro reached 45.0% (baselines max 11.0%) and LIBERO-OOD 63.9%. Adding a VLM planner lifted RoboDojo long-horizon success from 27.6% to 57.6%, pointing to a "GPT thinks, VPP2 acts" physical-AI paradigm. The company already runs routine operations with China Post and SF Express across 10+ logistics centers in 5 provinces. VPP2 is open-sourced.

Original post →

More from Embodied

Embodied channel →