AgentVLN: 3B VLM brain with skill library tops embodied navigation benchmarks

机器之心 · wechat · 2026-09-10

AgentVLN adopts a "VLM-as-Brain" design (accepted to ECCV 2026): a vision-language model handles task understanding and skill orchestration while modular skills handle mapping, obstacle avoidance and motion. It maps SLAM-computed 3D waypoints onto the camera image so the VLM picks paths visually, plus a self-correction mechanism and Query-Driven Perceptual CoT that actively queries depth sensing when uncertain.

Built on Qwen2.5-VL-3B, it runs in real time on Jetson edge devices, leads R2R-CE/RxR-CE benchmarks, and has been deployed on a quadruped and a humanoid robot.

Original post →

More from Embodied

Embodied channel →