Google Unveils Three Embodied AI Models, Solving Robotic 'Temporal Intelligence'
aigclink · x · 2026-07-31
Google has released three models for embodied AI, aiming to solve 'temporal intelligence'—allowing robots to know not only how to perform a task but also when it is completed.
- Gemini Robotics 2: A Vision-Language-Action (VLA) model controlling physical robot movements (closed-source).
- Gemini Robotics ER 2: An embodied reasoning model acting as the 'brain,' capable of watching video, planning multi-step tasks, and delegating execution to any VLA (closed-source).
- On-Device 2: Runs locally and can adapt to a new robot body within hours (open-source).
Key Breakthroughs in ER 2:
- Progress Classification: Categorizes video frames into five progress tiers with 57.4% accuracy.
- Precise Timing: Identifies the exact frame for key events (e.g., stopping coffee pouring) with 91.3% accuracy and an average error of just 0.96 seconds, achieving sub-second latency with low compute.
- Seamless Interaction: Integrates with the bidirectional Gemini Live API for smooth 'thinking while doing' and natively calls tools like Google Search.
- Multi-Robot Collaboration: Demonstrated task handovers between Apollo 2 and Franka F3 Duo, plus voice commands for Boston Dynamics Spot.
ER 2 is now available to developers via the Gemini API and AI Studio.
Related event: Google DeepMind Unveils Gemini Robotics 2(69 posts)→
More from Embodied
- AI Makes Call of Duty Heartbeat Sensor Real: Sensing the Physical World via RF — bilawalsidhu · 2026-08-01
- VLM run Integrates Gemini Robotics ER 2 into Orion 2 Visual Agent Harness — spillai · 2026-08-01
- Robot Goal-Reaching Test: Model Unexpectedly Breaks Into a Funny Dance — carlosdponx · 2026-08-01
- SAMRA Home Companion Robot: Multi-Axis Design Tailored for Children — tobowers · 2026-08-01
- Musk Calls Human Language 'Lossy Compression', Predicts Neuralink Will Replace Speech — r0ck3t23 · 2026-08-01
- Google DeepMind Unveils Gemini Robotics 2 Model — TheAIGRID · 2026-08-01