Mistral’s 8B Robostral Navigate uses one RGB camera and tops R2R-CE at 77.4%
mistralai · hf · 2026-07-24
What Mistral released
Robostral Navigate is an 8B vision-language model for robot navigation that uses only a single monocular RGB stream to predict the next waypoint in image space.
Why it matters
The model is designed to avoid depth sensors, multi-camera rigs, and prebuilt maps, making deployment easier across different robot embodiments.
Training recipe
- 2.4 million trajectories
- 350k simulated scenes
- A prefix-caching recipe that packs full episodes into single sequences
- Training tokens reduced by 22x, cutting training time from months to days
- A tree-based attention mask to prevent leakage from ground-truth actions
- Reinforcement learning to improve exploration and recovery
Results
- R2R-CE: 77.4% success rate, +10.5 points over the best monocular method and +5.3 points over the strongest depth- or multi-camera system
- RxR-CE: 75.1% success rate, beating all monocular baselines
Takeaway
Mistral is pushing robot navigation toward a cheaper, single-camera, cross-embodiment deployment model.
More from Embodied
- REK teases a samurai robot warm-up before tonight’s Tokyo fight night — cixliv · 2026-07-24
- SAGE generates simulation-ready 3D scenes and releases a 10k embodied-AI dataset — rsasaki0109 · 2026-07-24
- Tsinghua team turns two images into 3D mirror-illusion art with AutoMIA — 新智元 · 2026-07-24
- Black Forest Labs says FLUX 3 beats major multimodal rivals and powers robotics — Latent Space · 2026-07-24
- ReferTrack tracks language-specified targets with one camera and reaches 89.4% on EVT-Bench — tencent · 2026-07-24
- IGGT4D brings streaming 4D instance-grounded geometry to dynamic scene understanding — zhenjun_zhao · 2026-07-24