Grounded Action Model: Building Robot Foundation Models on 3D Grounding, Hits 61% on LIBERO-PRO
DJiafei · x · 2026-09-22
Researchers propose Grounded Action Model (GAM), a new paradigm building robot foundation models on pretrained 3D grounding rather than language (VLA) or video (WAM) backbones.
- Language/point/box prompts are converted into an object-centric representation mixing target visual features and metric geometry, fused with robot state history via a multi-stream transformer to predict action chunks.
- Can act autonomously or serve as a low-level controller under a high-level planner (Molmo2); on Franka, GAM+Molmo2 achieves 64.7% in-distribution and 49.8% OOD step completion.
- RoboTwin 2.0: 55.3% average success across 50 tasks (vs 52.0% Spatial Forcing), 47.6% under scene randomization (vs 30.4% Abot-M0), trained only on clean-scene demos.
- LIBERO-PRO: 61% average across 16 settings (vs 53% for π0.5); 88% vs 1% on target-switch tasks.
- Limitations: grounding errors and omitted context remain. Paper, code, and project page are public.
Related event: NUS Proposes GAM: 3D Grounding as Foundation for Robot Models(3 posts)→
More from Embodied
- RoboDawn: Tsinghua & Tencent Hunyuan drive robots with frozen VLM, 73.6% one-shot success — _akhaliq · 2026-09-22
- PragmaBot: robots learn online from real-world failures without retraining, RA-L/IROS 2026 — ChongZzZhang · 2026-09-22
- Humanoid demand could hit hundreds of millions; millions built by 2029, says Robostrategy — Rewkang · 2026-09-22
- Intrinsic open-sources Intrinsic Core: ROS-compatible building blocks for physical AI robotics — facontidavide · 2026-09-22
- MolmoAct2 gets out-of-the-box support for SO-101 arms and YAM, plus VR teleop — k7agar · 2026-09-22
- Quest-based robot data collection adds passthrough and headless modes, YAM teleop switching — k7agar · 2026-09-22