SJTU and Alibaba Introduce LA4VLA: Decoupling Language-Action to Boost Robot Policies

青稞AI · wechat · 2026-08-03

To address the issue of Vision-Language-Action (VLA) models over-relying on visual shortcuts and weakening language constraints, Shanghai Jiao Tong University and Alibaba introduced LA4VLA. This method temporarily removes visual input during the pre-training phase, allowing the model to focus exclusively on learning the correspondence between language and actions.

By decoupling language-action learning from visual grounding, the researchers constructed a 33K vision-agnostic dataset. Experiments demonstrate that this explicit language-action pre-training serves as an effective complementary signal to standard VLA training, significantly improving policy performance and robustness against visual perturbations in embodied robotics.

Original post →

More from Embodied

Embodied channel →