966 real navigation episodes, 24km: ECCV study reveals which visual encoder components matter for robots

chriswolfvision · x · 2026-09-12

A companion post to the ECCV paper (arXiv:2606.21216) by Steeve Janny, Leonid Antsfeld and Christian Wolf: 966 static point-goal navigation episodes over 24km in a real building, evaluating state-of-the-art visual encoders under realistic conditions. Key findings: encoder power stems largely from CV pre-training losses; heterogeneous multi-teacher distillation yields complementary skills; principled spatial bottlenecks produce interpretable affordance-linked features; and pre-training policies on privileged Lidar input before fine-tuning beats RGB-only training.

Related event: One scalar per patch suffices for real-world robot navigation, ECCV study shows(2 posts)→

Original post →

More from Embodied

Embodied channel →