Zero-Shot Generalist Agents Beat Trained Models in Vision-Language Navigation Without Scaffolding

机器之心 · wechat · 2026-08-27

Researchers from the University of Adelaide propose MIP (Minimal-Interface Probe), connecting general-purpose coding-agents (like Claude/Codex) directly to robots for Vision-and-Language Navigation (VLN). By removing traditional scaffolds like mapping, memory, and waypoint predictors—leaving only observe() and step()—the method achieved a 78% success rate on the R2R-CE benchmark, matching or exceeding industrial-scale specialized policies (human performance is 94%). Results show model reasoning is the core factor, and optional tools outperform forced integration. However, success rates drop to 26% on long-horizon tasks (RxR-CE), and real-world tests revealed a lack of proprioception (e.g., the body getting stuck while the camera passes through a door).

Original post →

More from Embodied

Embodied channel →