AgentVLN: a 3B VLM brain tops R2R-CE and RxR-CE in real time on Jetson

jiqizhixin · x · 2026-09-19

AgentVLN (ECCV 2026) introduces a "VLM-as-Brain" paradigm for vision-and-language navigation: Qwen2.5-VL-3B handles high-level decisions while a modular skill library (SLAM/perception/navigation) executes actions — no heavy 3D vision modules needed.

Original post →

More from Embodied

Embodied channel →