ABot-N1: A General Vision-Language Navigation Model

acvlab · hf · 2026-07-14

This paper proposes ABot-N1, aiming to build a general foundation model for vision-language navigation.

The authors argue that existing methods commonly map observations directly to actions, leading to coordinate drift, poor handling of long-tail semantics, and black-box inexplicability. To address this, ABot-N1 adopts a "slow-fast" architecture to decouple cognition from control:

The paper claims this approach is more robust across simulated and real-world benchmarks, achieving significant gains in city-scale navigation: POI arrival improved by 35.0% to reach 77.3%; SR reached 95.4%/92.9% in complex indoor/outdoor scenes. Furthermore, the authors open-sourced new Point-Goal / POI-Goal benchmarks.

Original post →

More from Embodied

Embodied channel →