SJTU Team Proposes GSR to Tackle Instruction Generalization in VLA Models
机器之心 · wechat · 2026-08-06
Current Vision-Language-Action (VLA) models suffer from severe instruction generalization flaws, where simply paraphrasing a task command causes action success rates to plummet. A joint team from SJTU and Boundless Power found that equivalent wording introduces feature shifts during multimodal fusion, which are then amplified by downstream action modules.
To solve this, they proposed GSR (Grounded Semantic Re-Binding): using a frozen T5 encoder to independently extract stable task semantics, re-injecting them into the feature flow based on different VLA architectures, and retraining the action expert. Experiments show that without training on paraphrased instructions, this method significantly boosts generalization in lightweight models (e.g., SmolVLA jumps from 4.47% to 49.12%) and provides gains even for large pre-trained models.
More from Embodied
- Indie Dev Uses AI to Design Circuit Boards, Aims to Crack Production and Sell Hardware — pramodk73 · 2026-08-06
- US Robot Ban Hits Startups: Requires Over 65% Domestic BOM Sourcing — mattfreed · 2026-08-06
- BridgeVLA++ Boosts 3D Robotic Manipulation with Spatio-Temporal Memory Architecture — Peiyan Li · 2026-08-06
- Nori L3 Dual-Arm Home Robot Launched at $1,688 — CyberRobooo · 2026-08-06
- PowerBot's Multi-Sport AI Coach Robot Raises Over $4M on Kickstarter — 创业邦 · 2026-08-06
- Beijing Deploys 72 Park Robots for Patrolling, Cleaning, and Pest Control — pstAsiatech · 2026-08-06