Grounded Action Model: Building robot foundation models on 3D grounding instead of VLAs
DJiafei · x · 2026-09-23
Researchers introduce Grounded Action Model (GAM), a new paradigm for robot manipulation foundation models: instead of building on language or video models (VLAs/WAMs), GAM is constructed on top of a pretrained 3D grounding model. The core idea is "ground first, then learn to act," claiming 3D grounding yields more capable and generalizable manipulation. Details are shared in a thread.
Related event: GAM: Grounded Action Model Builds Robot Foundation Models on 3D Grounding(6 posts)→
More from Embodied
- LIBERO experiments show stronger reasoning in GPT-6 variants means better robot manipulation — YuXiang_IRVL · 2026-09-23
- eidon-ai releases tracker-pov robotics video dataset, 10K-100K entries — eidon-ai · 2026-09-23
- Dev turns a 90s TV into an AI TV with Raspberry Pi, mic, camera and GPT Realtime — OpenAIDevs · 2026-09-23
- Robot arm tests rank GPT-6 variants: better reasoning means better manipulation at 30x the cost — YuXiang_IRVL · 2026-09-23
- Steering VLA robot models at inference time lifts grasp rate from 14% to 74% without retraining — drmapavone · 2026-09-23
- Frontier AI models can control robots to follow harmful requests, NBC News reports — ycombinator · 2026-09-23