Peking Univ. & MSRA Introduce BCP: Boosting Robot Success Rates via Autonomous Replanning

jiqizhixin · x · 2026-08-26

Peking University and Microsoft Research Asia introduced BCP (Bernoulli-Continuation Policy), a method enabling robots to autonomously decide whether to continue executing an action chunk or stop and replan.

BCP freezes the base VLA model and trains a tiny 16.4M-parameter head to model execution horizons as a sequence of Bernoulli decisions. It optimizes via GRPO with a reward function that penalizes excessive VLA calls. Results show LingBot-VLA's success rate on RoboTwin 2.0 tasks jumped from 89.88% to 93.94% (SOTA among VLA methods). Real-world robot mug-hanging success soared from 44% to 84%. Despite more frequent replanning, total runtime decreased due to higher accuracy and reduced wasted motion.

Original post →

More from Embodied

Embodied channel →