BIND Action Head Ties Robot Actions to 2D Image Features for Data-Efficient Policies

yuewang314 · x · 2026-10-03

Researchers highlight a paradox in robot learning: image encoders are spatially smart — features embed semantics, geometry, and are multiview-consistent — but robot policies built on them are spatially dumb, needing hundreds of demos for simple pick-and-place and breaking under small camera shifts.

Their answer is BIND, a new action head that binds each candidate robot action to its projected 2D image feature(s), effectively telling the network "choosing this action moves the EEF to this image feature."

Result: far more data-efficient policies that are robust to OOD viewpoints and object positions, using only RGB input with no 3D sensors, outputting dense EEF trajectories.

Original post →

More from Embodied

Embodied channel →