ECCV paper: one scalar per patch from pre-trained ViTs enables fast real-world robot navigation
chriswolfvision · x · 2026-09-12
Christian Wolf's team presents an ECCV study on real-world robot navigation: visual encoders distilled from heterogeneous teachers can be bottlenecked to just one scalar per patch via attention projection, yet still support fast moving navigation in a real building. Policies are pre-trained with privileged Lidar input and then fine-tuned to RGB-only. An interpretable affordance-linked structure emerges. A large-scale 966-episode / 24km real-robot evaluation is covered in a companion post.
More from Embodied
- Weekend robot fights at autonomous labs: train your bot by chatting with an AI coach — dee_hw · 2026-09-12
- Flying humanoid robot installs AC on 30th floor in China, remote operator replaces rope climbers — TansuYegen · 2026-09-12
- 11 Years of Humanoid Robotics: From DARPA Falls to 2,056-Robot Games — TinfoilTricorn · 2026-09-12
- PokeBot CTO predicts zero-shot generalization on 80% of household tasks within a year — suchenzang · 2026-09-12
- GPT-6 Astra rumored to use recurrent-depth latent reasoning, sparking AI safety alarms — DJiafei · 2026-09-12
- Robot drags a 50kg rice bag needlessly — how do we measure robot "stupidity"? — ducha_aiki · 2026-09-12