Latent Interface Training boosts VLA generalization, lifting LIBERO-Plus success up to 11 points
DJiafei · x · 2026-09-11
- Problem: backgrounds, camera angles, and lighting in training data can correlate with demonstrated actions, letting VLA/WAM action experts learn visual shortcuts that break in new scenes.
- Method — LIT (Latent Interface Training), two stages:
- Learn to act first: train the action expert without images, conditioned on language, robot state, and each action chunk's terminal end-effector pose;
- Learn to use vision: initialize from stage 1 and route visual conditioning through a latent interface supervised to reconstruct the same terminal pose, preserving goal-relevant spatial information.
- Results: average LIBERO-Plus success improves across four architectures — π0.5 68.97%→79.67%, MolmoAct2 63.62%→71.92%, FAST-WAM 51.44%→60.63%, ImageWAM 83.02%→86.89%. In a real-world test, MolmoAct2 learned 3 tasks from 300 demonstrations, with new-lighting success rising 53.3%→70.0%, top-camera-only 30.0%→46.7%, distractors 50.0%→63.3%, in-distribution 74.7%→88.0%.
Related event: LIT Breaks Vision-Action Shortcuts in Robot VLA Models(4 posts)→
More from Embodied
- Skild hits $100M revenue run rate; XPENG starts IRON humanoid production line — xmercury_one · 2026-09-11
- Gary Marcus: AI is strong in constrained domains, weak in the open physical world — GaryMarcus · 2026-09-11
- Latent Interface Training breaks vision-action shortcuts in VLA and WAM training — DJiafei · 2026-09-11
- Robot nearly sticks the landing: training progress but leg lift still fails — hbouammar · 2026-09-11
- Robot's kick attempt goes from near-fall to almost sticking the landing — 0xSammy · 2026-09-11
- Robots can dance, do martial arts and finish triathlons — but still can't wash dishes — tinyfool · 2026-09-11