IROS 2026 Best Paper: Feeding 3D Features Directly into VLM Lifts Navigation Success to 74.2%
qinzytech · x · 2026-10-09
SoftNav, the IROS 2026 Best Paper, passes learned 3D scene features into a vision-language model as "soft tokens" to guide navigation, instead of converting scenes to text.
- On HM3D-OVON's seen split, with the same 3D encoder and VLM, direct features raised navigation success from 66.7% (text interface) to 74.2%
- On the unseen split results diverged: SoftNav reached more goals (66.7% vs 63.3%) but scored lower on path-length-weighted success (25.7% vs 28.3%), taking longer routes on solved episodes
- A richer text baseline statistically matched SoftNav's success rate
- Training used just 1,187 examples and 17M trainable parameters, with base models frozen
The author asks: have you compared text vs direct features with fixed base models, and did the gap show in goal-reaching or route efficiency?
More from Embodied
- Text2Sim turns text prompts into editable physics-grounded videos and sim setups — erwincoumans · 2026-10-09
- Singapore's first giant robot fight flopped because teleop policies weren't built for combat, researcher argues — carlosdponx · 2026-10-09
- PredActor unifies humanoid motion generation and control in one diffusion policy at 50Hz on Unitree G1 — carlosdponx · 2026-10-09
- Meta Glasses US Open ad goes viral: your phone might just stay in your pocket — armand_ruiz · 2026-10-09
- Percentile-controlled diffusion generates calibrated corner-case scenarios for AV safety testing — Jiaxi Liu · 2026-10-09
- Intel's John Healy: most enterprise robotics value sits in the installed base — ryanshrout · 2026-10-09