LIT trains VLAs to act before seeing: two-stage method fights visual shortcuts

DJiafei · x · 2026-09-11

Standard VLAs/WAMs train action experts directly on visual representations, but backgrounds, camera angles, or lighting in training data can correlate with demonstrated actions — the policy learns these visual shortcuts and fails when correlations shift in new scenes.

LIT (Latent Interface Training) restructures training into two stages:

This constrains visual reliance while preserving goal-relevant spatial info, improving robustness out of distribution.

Related event: LIT Breaks Vision-Action Shortcuts in Robot VLA Models(4 posts)→

Original post →

More from Embodied

Embodied channel →