LIT trains VLAs to act before seeing: two-stage method fights visual shortcuts
DJiafei · x · 2026-09-11
Standard VLAs/WAMs train action experts directly on visual representations, but backgrounds, camera angles, or lighting in training data can correlate with demonstrated actions — the policy learns these visual shortcuts and fails when correlations shift in new scenes.
LIT (Latent Interface Training) restructures training into two stages:
- Learn to act first: train the action expert without images, conditioned on language, robot state, and each demonstrated action chunk's terminal end-effector pose.
- Learn to use vision: initialize from Stage 1, then jointly train with visual conditioning routed exclusively through a latent interface, supervised to reconstruct the same terminal pose — encouraging it to retain goal-relevant spatial information.
This constrains visual reliance while preserving goal-relevant spatial info, improving robustness out of distribution.
Related event: LIT Breaks Vision-Action Shortcuts in Robot VLA Models(4 posts)→
More from Embodied
- Skild hits $100M revenue run rate; XPENG starts IRON humanoid production line — xmercury_one · 2026-09-11
- Gary Marcus: AI is strong in constrained domains, weak in the open physical world — GaryMarcus · 2026-09-11
- Latent Interface Training breaks vision-action shortcuts in VLA and WAM training — DJiafei · 2026-09-11
- Robot nearly sticks the landing: training progress but leg lift still fails — hbouammar · 2026-09-11
- Robot's kick attempt goes from near-fall to almost sticking the landing — 0xSammy · 2026-09-11
- Robots can dance, do martial arts and finish triathlons — but still can't wash dishes — tinyfool · 2026-09-11