SJTU and Alibaba Introduce LA4VLA: Decoupling Language-Action to Boost Robot Policies
青稞AI · wechat · 2026-08-03
To address the issue of Vision-Language-Action (VLA) models over-relying on visual shortcuts and weakening language constraints, Shanghai Jiao Tong University and Alibaba introduced LA4VLA. This method temporarily removes visual input during the pre-training phase, allowing the model to focus exclusively on learning the correspondence between language and actions.
By decoupling language-action learning from visual grounding, the researchers constructed a 33K vision-agnostic dataset. Experiments demonstrate that this explicit language-action pre-training serves as an effective complementary signal to standard VLA training, significantly improving policy performance and robustness against visual perturbations in embodied robotics.
More from Embodied
- Tacta Systems emerges from stealth with $75M, launches dexterous robot hand TactaBot — ZeYanjie · 2026-08-03
- China's House Cleaning Robots Cost ~$17 for 3 Hours, Include Human Supervision — MarwaEldiwiny · 2026-08-03
- OpenArm: Open-Source 7-DOF Humanoid Arm for Physical AI Research — tom_doerr · 2026-08-03
- Humanoid Robots in Logistics: RobotEra M7 Sorts 1,200 Packages Per Hour — CyberRobooo · 2026-08-02
- New AI Learns Parkour From Just 30 Seconds Of Video — Two Minute Papers · 2026-08-02
- US FCC Bans Imports of Foreign Humanoid Robots Over Security Fears — heatherknight · 2026-08-02