966 real navigation episodes, 24km: ECCV study reveals which visual encoder components matter for robots
chriswolfvision · x · 2026-09-12
A companion post to the ECCV paper (arXiv:2606.21216) by Steeve Janny, Leonid Antsfeld and Christian Wolf: 966 static point-goal navigation episodes over 24km in a real building, evaluating state-of-the-art visual encoders under realistic conditions. Key findings: encoder power stems largely from CV pre-training losses; heterogeneous multi-teacher distillation yields complementary skills; principled spatial bottlenecks produce interpretable affordance-linked features; and pre-training policies on privileged Lidar input before fine-tuning beats RGB-only training.
More from Embodied
- Awesome AI Hardware adds 8 vetted open-source AI×hardware projects — sujingshen · 2026-09-12
- Open-source digital koi pond doubles as a voice recorder for under $60 — TinfoilTricorn · 2026-09-12
- Adam Dorr asks if US should ban humanoid robot exports to all foreign countries — adam_dorr · 2026-09-12
- GPT-6 Astra shows 'step change' in spatial reasoning, solving 7/100 robot tasks vs zero for rival — The Decoder · 2026-09-12
- China's snake-like grid robot draws power from live wires, runs 24/7 inspections — prasanna_says · 2026-09-12
- Headsets hit 100-200g yet the criticism never stops: a dev's take on the VR doom loop — Scobleizer · 2026-09-12