ECCV paper: one scalar per patch from pre-trained ViTs enables fast real-world robot navigation

chriswolfvision · x · 2026-09-12

Christian Wolf's team presents an ECCV study on real-world robot navigation: visual encoders distilled from heterogeneous teachers can be bottlenecked to just one scalar per patch via attention projection, yet still support fast moving navigation in a real building. Policies are pre-trained with privileged Lidar input and then fine-tuned to RGB-only. An interpretable affordance-linked structure emerges. A large-scale 966-episode / 24km real-robot evaluation is covered in a companion post.

Related event: One scalar per patch suffices for real-world robot navigation, ECCV study shows(2 posts)→

Original post →

More from Embodied

Embodied channel →