LIT lifts real-robot success rates 13-17 points by breaking vision-action shortcuts
DJiafei · x · 2026-09-11
- Paper from NUS Magic Lab et al.: VLAs and WAMs learn vision-action shortcuts (lighting, background, camera view correlated with actions), hurting OOD generalization.
- LIT trains in two stages around a shared spatial goal: first train the action expert without images (language, robot state, target end-effector pose), then route vision through latent tokens supervised to predict the same pose.
- Real-robot results on a YAM dual-arm platform: in-distribution 88.0% vs 74.7% baseline (+13.3pp); new lighting +16.7pp, new camera +16.7pp, distractors +13.3pp. Only LIT redirects when the instructed goal changes.
- Sim results on LIBERO-Plus improve average success across four architectures including π0.5. Paper, code, and project page are open.
Related event: LIT Breaks Vision-Action Shortcuts in Robot VLA Models(4 posts)→
More from Embodied
- Autonomous Lamp open-source companion robot adds dust, air, temperature and humidity sensing — dee_hw · 2026-09-11
- Cybertruck is Tesla's rolling engineering lab: 48V, steer-by-wire, Etherloop — XFreeze · 2026-09-11
- Analysts break down Apple's fall keynote: foldable iPhone and the Ternus era — BenBajarin · 2026-09-11
- Autonomous WorkPod bundles solar, Starlink, and dual RTX 5090s into a $20,900 backyard AI datacenter — dee_hw · 2026-09-11
- First-time PCB designer vibe-codes a working e-ink dev board with Claude, pays €130 to fab it — a6m--zero · 2026-09-11
- DeepMind open-sources flybody: a full fruit fly simulated joint-by-joint in MuJoCo, published in Nature — ZeroStateReflex · 2026-09-11