Grounded Action Model: Building robot foundation models on 3D grounding instead of VLAs

DJiafei · x · 2026-09-23

Researchers introduce Grounded Action Model (GAM), a new paradigm for robot manipulation foundation models: instead of building on language or video models (VLAs/WAMs), GAM is constructed on top of a pretrained 3D grounding model. The core idea is "ground first, then learn to act," claiming 3D grounding yields more capable and generalizable manipulation. Details are shared in a thread.

Related event: GAM: Grounded Action Model Builds Robot Foundation Models on 3D Grounding(6 posts)→

Original post →

More from Embodied

Embodied channel →