Harness VLA: Boosting Frozen Robot Models Without Retraining

jiqizhixin · x · 2026-07-31

Researchers from Tsinghua University, Striding AI, and Purdue University introduced Harness VLA, a framework designed to help robots reliably follow spoken commands even when environments change.

Instead of retraining Vision-Language-Action (VLA) models from scratch, Harness VLA wraps a frozen VLA with a smart planner and a small library of simple commands (MOVE TO, ROTATE, GRAB). The planner uses memory from past successes and failures to decide whether to invoke the complex VLA for tricky contact-heavy tasks or rely on simpler analytic steps.

The approach outperforms the strongest baselines by 38.6 and 25.4 percentage points on the LIBERO-Pro and RoboCasa365 benchmarks respectively, and achieves 58.4% on RoboTwin C2R, proving that frozen VLAs can be pushed significantly further.

Original post →

More from Embodied

Embodied channel →