GAM repurposes a geometric foundation model as robot manipulation policy, moving beyond 2D VLA

rsasaki0109 · x · 2026-09-02

The author proposes Geometric Action Model (GAM), a language-conditioned manipulation policy that reuses a pretrained Geometric Foundation Model (GFM) as a shared substrate for perception, temporal prediction, and action decoding.

GAM splits the GFM at an intermediate layer: shallow layers act as an observation encoder while deeper parts handle causal future prediction and action decoding. The argument: existing VLA and video world-action models inherit semantic/temporal priors but operate on 2D frames or 2D-derived latents, leaving the geometry needed for contact-rich manipulation implicit. Paper and code are linked on GitHub.

Original post →

More from Embodied

Embodied channel →