AgentVLN: a 3B VLM brain tops R2R-CE and RxR-CE in real time on Jetson
jiqizhixin · x · 2026-09-19
AgentVLN (ECCV 2026) introduces a "VLM-as-Brain" paradigm for vision-and-language navigation: Qwen2.5-VL-3B handles high-level decisions while a modular skill library (SLAM/perception/navigation) executes actions — no heavy 3D vision modules needed.
- A QD-PCoT mechanism generates fine-grained exploration under occlusion, blind spots and trajectory drift, curbing long-horizon error accumulation without extra parameters
- Leads the mainstream R2R-CE and RxR-CE benchmarks while running in real time on Jetson edge devices
More from Embodied
- Y Combinator-backed Proception to debut new humanoid hand tech at IROS 2026 — chris_j_paxton · 2026-09-19
- Do VLMs Plagiarize? UC Berkeley's Jitendra Malik Weighs In on Credit Assignment in the Astra Demo Debate — JitendraMalikCV · 2026-09-19
- micro1: frontier models controlling real hardware is a pressing AI safety problem — Exp_Mark · 2026-09-19
- Simulated Claude arms spill toxic liquids and use harmful force in micro1 safety tests — Exp_Mark · 2026-09-19
- Robotics has hit its tipping point, says analyst firm after visiting a leading lab — citrini · 2026-09-19
- Robot Control Explained for ML People: PID, FK/IK, and Why ACT Outputs Joint Angles — kamathsblog · 2026-09-19