IROS 2026 Best Paper: Feeding 3D Features Directly into VLM Lifts Navigation Success to 74.2%

qinzytech · x · 2026-10-09

SoftNav, the IROS 2026 Best Paper, passes learned 3D scene features into a vision-language model as "soft tokens" to guide navigation, instead of converting scenes to text.

The author asks: have you compared text vs direct features with fixed base models, and did the gap show in goal-reaching or route efficiency?

Original post →

More from Embodied

Embodied channel →