Gemini Robotics Model Accurately Splits 9-Minute Excavator Video into 40+ Clips

DynamicWebPaige · x · 2026-07-30

Google's Gemini Robotics model demonstrates advanced video understanding capabilities. Tested on a 9-minute video of an excavator loading trucks, the model successfully segmented the footage into over 40 individual clips, accurately annotating each dirt pick-up and dump action along with its location.

Furthermore, the model can automatically tally the frequency of specific actions, track the types of equipment entering a zone, and aggregate sensor readings, achieving multi-dimensional structured analysis of complex long videos.

Original post →

More from Embodied

Embodied channel →