ABot-N1: A General Vision-Language Navigation Model
acvlab · hf · 2026-07-14
This paper proposes ABot-N1, aiming to build a general foundation model for vision-language navigation.
The authors argue that existing methods commonly map observations directly to actions, leading to coordinate drift, poor handling of long-tail semantics, and black-box inexplicability. To address this, ABot-N1 adopts a "slow-fast" architecture to decouple cognition from control:
- Slow vision-language reasoner: First performs explicit Chain-of-Thought reasoning and generates pixel-level target points.
- Pixel anchors: These image-space anchors act as a universal interface for tasks like point-goal, object-goal, poi-goal, instruction following, and person following.
- Fast action expert: Combines textual cues and pixel guidance to output continuous waypoints at the native control frequency.
The paper claims this approach is more robust across simulated and real-world benchmarks, achieving significant gains in city-scale navigation: POI arrival improved by 35.0% to reach 77.3%; SR reached 95.4%/92.9% in complex indoor/outdoor scenes. Furthermore, the authors open-sourced new Point-Goal / POI-Goal benchmarks.
More from Embodied
- Tesla expands Robotaxi rides to seven areas, including new Orlando and Tampa zones — elonmusk · 2026-07-22
- Hands-on robotics workshop on Saturday may be the last in-person session before August — StewartalsopIII · 2026-07-22
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21
- Gritt says an 8-person crew now installs 3,000 to 4,000 solar panels a day — HaktanSuren · 2026-07-21
- A helium-powered flying robot whale aims to be a quiet companion pet — chris_j_paxton · 2026-07-21