PerturBot Breaks Shortcut Priors in Vision-Language-Action Models With Perturbative Training

Mingyu Liu · hf · 2026-10-06

Georgia Tech researchers identify "modality shortcuts" in VLA policies: models exploit surface regularities in demonstrations—ignoring evidence like objects displaced near the wrist camera or changed verbs—rather than true task-relevant cues.

PerturBot counters this with task-preserving wrist-view perturbations, decision-relevant caption-enriched instructions, and relabeled random/failed trajectory segments, all without changing inference. It complements scaling rather than replacing it.

They also propose GroundingFscore, an offline metric diagnosing how severely a policy relies on shortcuts—revealing whether scaling is healthy beyond raw task success rates.

Original post →

More from Embodied

Embodied channel →