SJTU Team Proposes GSR to Tackle Instruction Generalization in VLA Models

机器之心 · wechat · 2026-08-06

Current Vision-Language-Action (VLA) models suffer from severe instruction generalization flaws, where simply paraphrasing a task command causes action success rates to plummet. A joint team from SJTU and Boundless Power found that equivalent wording introduces feature shifts during multimodal fusion, which are then amplified by downstream action modules.

To solve this, they proposed GSR (Grounded Semantic Re-Binding): using a frozen T5 encoder to independently extract stable task semantics, re-injecting them into the feature flow based on different VLA architectures, and retraining the action expert. Experiments show that without training on paraphrased instructions, this method significantly boosts generalization in lightweight models (e.g., SmolVLA jumps from 4.47% to 49.12%) and provides gains even for large pre-trained models.

Original post →

More from Embodied

Embodied channel →