DeepSeek's New V4 Flash Vision Model Integrated into Humanoid Robot for Navigation Task

zizhpan · x · 2026-08-22

DeepSeek released the V4 Flash Vision model, combining the agent/reasoning capabilities of V4 Flash with visual understanding. According to benchmarks, it approaches Anthropic's Opus 4.8 on multimodal agent tasks.

The author integrated this model into a humanoid robot named "RalCox" and gave it a goal: "Safely approach the person using the laptop." The robot captured images via its onboard camera, sent them to DeepSeek, and successfully executed the task. The screenshots show the live camera view and DeepSeek's reasoning process about what the robot sees.

Original post →

More from Embodied

Embodied channel →