DeepSeek's New V4 Flash Vision Model Integrated into Humanoid Robot for Navigation Task
zizhpan · x · 2026-08-22
DeepSeek released the V4 Flash Vision model, combining the agent/reasoning capabilities of V4 Flash with visual understanding. According to benchmarks, it approaches Anthropic's Opus 4.8 on multimodal agent tasks.
The author integrated this model into a humanoid robot named "RalCox" and gave it a goal: "Safely approach the person using the laptop." The robot captured images via its onboard camera, sent them to DeepSeek, and successfully executed the task. The screenshots show the live camera view and DeepSeek's reasoning process about what the robot sees.
More from Embodied
- Blind user praises Tesla Cybercab for Braille support and lack of discrimination — EricETesla · 2026-08-22
- NVIDIA's ADEPT Pre-trains Dexterity, Cuts Robot Task Learning Cost by 3x — KyleMorgenstein · 2026-08-22
- Nobel laureate Szostak discusses AI robot self-evolution hypothesis — danfaggella · 2026-08-22
- My custom humanoid robot parts supplier is in China — 'Not sure the US can do better' — TinfoilTricorn · 2026-08-22
- Two Ludi agents talking to each other after a tough demo day — Kangwook_Lee · 2026-08-22
- Engineer develops hovering umbrella using drone technology — mtizard · 2026-08-22