PerturBot Breaks Shortcut Priors in Vision-Language-Action Models With Perturbative Training
Mingyu Liu · hf · 2026-10-06
Georgia Tech researchers identify "modality shortcuts" in VLA policies: models exploit surface regularities in demonstrations—ignoring evidence like objects displaced near the wrist camera or changed verbs—rather than true task-relevant cues.
PerturBot counters this with task-preserving wrist-view perturbations, decision-relevant caption-enriched instructions, and relabeled random/failed trajectory segments, all without changing inference. It complements scaling rather than replacing it.
They also propose GroundingFscore, an offline metric diagnosing how severely a policy relies on shortcuts—revealing whether scaling is healthy beyond raw task success rates.
More from Embodied
- Prompt-Driven Autonomous Drone Flies Fully Unsupervised After Start — AryHHAry · 2026-10-06
- Survey: 65% of Japanese seniors prefer robot-assisted nursing homes, willing to pay 8% more — HealthcareLdr · 2026-10-06
- Tianjin University unveils 3-gram hair-concealable non-invasive brain-computer interface — Dr_Singularity · 2026-10-06
- FlashDexRetarget: one RL policy retargets hand demos at 90% success, ~100x less compute — KyleMorgenstein · 2026-10-06
- Frankenduck robot rebuilt as fully parametric CAD parts editable via Python in STEVE studio — TinfoilTricorn · 2026-10-06
- PaXini's PX-Footrix plantar tactile soles arrive as robotics enters 'tactile sensing season' — LexiLove · 2026-10-06