DeepMind's EXIMO shows the harness-eating loop: VLM scaffold distilled into VLA weights for RL
m_wulfmeier · x · 2026-09-07
A Google DeepMind researcher highlights EXIMO, a student researcher project, as a clean example of the create-and-eat-harness tension in physical AI. EXIMO wraps Gemini Robotics (a VLA) with Gemini (a VLM) for interpretable hierarchical control, then distills the orchestrated behavior into the VLA weights via simple imitation — enabling RL optimization of the whole model beyond what the scaffold could hand-design, readying the next generation of harnesses.
More from Embodied
- Figure's Index creator network processes 30 min of video per second; $15M paid to data creators — adcock_brett · 2026-09-08
- Sim2real nails it first try: a new film production pipeline via robot simulation — alexcovo_eth · 2026-09-08
- Robotics startup Action Intelligence unveils Continuo, debuts at ECCV Sept 10-12 — thetripathi58 · 2026-09-07
- Embodied AI shifts from leaderboards to deployment: Annu Intelligence cuts robot rollout costs by 80% — 智东西 · 2026-09-07
- Unitree's UnifoLM-X2-1.0 world model runs fully autonomous humanoid fights in real time — Distinct-Question-16 · 2026-09-07
- World Labs unveils Atlas world model: native text/image/video/3D, 1-min 1440p camera-controlled video — thione · 2026-09-07