SuperNav: ZJU's Agent Harness Lets MLLMs Navigate Any Scene Without Fine-tuning
zju3dv · hf · 2026-10-09
ZJU's zju3dv lab unveiled SuperNav, an agentic navigation system that equips a pretrained MLLM with a navigation harness instead of fine-tuning it. The MLLM handles request interpretation, scene understanding, and decisions, while Navigation Skills tools handle motion execution via a unified visual-point interface. SuperNav beats four baselines on instance-level, multi-object, and demand-driven tasks, with category-level evaluation on HM3D and real quadruped robot deployment.
More from Embodied
- Scoble rides Tesla Cybercab: exceeds expectations, plans to buy one to rent out — Scobleizer · 2026-10-09
- Egocentric data collection is all the rage now, says researcher citing an ahead-of-its-time project — gragtah · 2026-10-09
- DreamTrue robot world model cuts interaction defect rate from 48.12% to 6.25%, tops AgiBot challenge — Junyan Li · 2026-10-09
- NVIDIA's USDCraft uses LLM-written programs to build simulation-ready articulated 3D assets — nvidia · 2026-10-09
- SimpleICL Defines Robot In-Context Learning with a Low-Cost Open Recipe — Minxing Li · 2026-10-09
- UBTECH inks FAW-Volkswagen deal for factory humanoids, targets 10,000 robots a year — CyberRobooo · 2026-10-09