AgentVLN: 3B VLM brain with skill library tops embodied navigation benchmarks
机器之心 · wechat · 2026-09-10
AgentVLN adopts a "VLM-as-Brain" design (accepted to ECCV 2026): a vision-language model handles task understanding and skill orchestration while modular skills handle mapping, obstacle avoidance and motion. It maps SLAM-computed 3D waypoints onto the camera image so the VLM picks paths visually, plus a self-correction mechanism and Query-Driven Perceptual CoT that actively queries depth sensing when uncertain.
Built on Qwen2.5-VL-3B, it runs in real time on Jetson edge devices, leads R2R-CE/RxR-CE benchmarks, and has been deployed on a quadruped and a humanoid robot.
More from Embodied
- NTU spin-off Ropedia launches HOMIE Gen 2 wearable system to train robots from human experience — liuziwei7 · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- Replit Agent in a robot builds and publishes websites autonomously via MCP — amasad · 2026-09-11
- TARS Robotics unveils embodied foundation model AWE: 15+ tasks, one model, zero retraining — heyshrutimishra · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11
- Working with an ESP32 device using Copilot CLI and Astra — DanWahlin · 2026-09-11