Mistral’s 8B Robostral Navigate uses one RGB camera and tops R2R-CE at 77.4%
mistralai · hf · 2026-07-24
What Mistral released
Robostral Navigate is an 8B vision-language model for robot navigation that uses only a single monocular RGB stream to predict the next waypoint in image space.
Why it matters
The model is designed to avoid depth sensors, multi-camera rigs, and prebuilt maps, making deployment easier across different robot embodiments.
Training recipe
- 2.4 million trajectories
- 350k simulated scenes
- A prefix-caching recipe that packs full episodes into single sequences
- Training tokens reduced by 22x, cutting training time from months to days
- A tree-based attention mask to prevent leakage from ground-truth actions
- Reinforcement learning to improve exploration and recovery
Results
- R2R-CE: 77.4% success rate, +10.5 points over the best monocular method and +5.3 points over the strongest depth- or multi-camera system
- RxR-CE: 75.1% success rate, beating all monocular baselines
Takeaway
Mistral is pushing robot navigation toward a cheaper, single-camera, cross-embodiment deployment model.
More from Embodied
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11
- MKBHD goes hands-on with the first folding iPhone; $2,000 48MP selfie cam mocked — alexmacgregor__ · 2026-09-11
- NTU spin-off Ropedia launches HOMIE Gen 2 wearable system to train robots from human experience — liuziwei7 · 2026-09-11