NUS releases LIT to break vision-action shortcuts in robot foundation models

NationalUniversityofSingapore · hf · 2026-09-14

NUS introduces LIT (Latent Interface Training), which improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information — breaking the vision-action shortcut in robotics foundation models.

Original post →

More from Embodied

Embodied channel →