Zero-Shot Generalist Agents Beat Trained Models in Vision-Language Navigation Without Scaffolding
机器之心 · wechat · 2026-08-27
Researchers from the University of Adelaide propose MIP (Minimal-Interface Probe), connecting general-purpose coding-agents (like Claude/Codex) directly to robots for Vision-and-Language Navigation (VLN). By removing traditional scaffolds like mapping, memory, and waypoint predictors—leaving only observe() and step()—the method achieved a 78% success rate on the R2R-CE benchmark, matching or exceeding industrial-scale specialized policies (human performance is 94%). Results show model reasoning is the core factor, and optional tools outperform forced integration. However, success rates drop to 26% on long-horizon tasks (RxR-CE), and real-world tests revealed a lack of proprioception (e.g., the body getting stuck while the camera passes through a door).
More from Embodied
- Anti-Surveillance Wearables: IR Blocking and Medical Masks May Become Common — TinfoilTricorn · 2026-08-27
- Hugging Face unveils new open source robot; co-founder Thom Wolf wants one — Thom_Wolf · 2026-08-27
- Agent Opens Bambu Handy App on Phone to Reprint Job — haydendevs · 2026-08-27
- Empatica's Parkinson's Monitoring Platform Receives FDA Clearance — RosalindPicard · 2026-08-27
- Pollen Robotics Releases Microduck Quadruped Robot — robotswantdata · 2026-08-27
- X-Square's WALL-SS World Model Enables Reliable Transfer from Virtual to Physical Robot Tasks — APPSO · 2026-08-27