Grounded Action Model builds robot foundation models on pretrained 3D grounding instead of VLMs

DJiafei · x · 2026-09-22

A new paradigm called Grounded Action Model (GAM) challenges the assumption that language or video models are the right foundation for robotics.

Inspired by developmental psychology (Hespos & Spelke's work on conceptual precursors to language), GAM argues robots should ground objects in 3D first, then learn to act — mirroring how infants grasp before they speak.

Key details:

Paper, code, and project page are released.

Related event: NUS Proposes GAM: 3D Grounding as Foundation for Robot Models(3 posts)→

Original post →

More from Embodied

Embodied channel →