Single RGB Camera Embodied Navigation Model Hits SOTA
sophiamyang · x · 2026-07-08
Achieving SOTA on the R2R-CE validation set (79.4% success in seen scenes, 76.6% in unseen), this model relies solely on a single RGB camera without LiDAR or depth sensors. The 8B model is trained entirely in simulation and runs across wheeled, legged, and aerial robots of varying sizes, while remaining robust to different camera intrinsics.
Related event: Mistral Unveils Robostral, an 8B Embodied Navigation Model(12 posts)→
More from Embodied
- Openloong shows a wheeled humanoid robot autonomously hauling trash bins in Shanghai — CyberRobooo · 2026-07-21
- Open-AoE opens 2,000 hours of egocentric manipulation video for robot learning — inclusionAI · 2026-07-21
- Blender depth maps drive an LTX-2.3 IC-LoRA video workflow in ComfyUI — waterarttrkgl · 2026-07-21
- AMD shows Ryzen AI Halo as a 100% local AI platform for on-device workflows — Sam Witteveen · 2026-07-21
- Halliday’s second-gen smart glasses fix the display problem from the original model — The Verge AI · 2026-07-21
- Halliday’s G2 smart glasses summarize meetings without using a camera — Wired AI · 2026-07-21