GAM repurposes a geometric foundation model as robot manipulation policy, moving beyond 2D VLA
rsasaki0109 · x · 2026-09-02
The author proposes Geometric Action Model (GAM), a language-conditioned manipulation policy that reuses a pretrained Geometric Foundation Model (GFM) as a shared substrate for perception, temporal prediction, and action decoding.
GAM splits the GFM at an intermediate layer: shallow layers act as an observation encoder while deeper parts handle causal future prediction and action decoding. The argument: existing VLA and video world-action models inherit semantic/temporal priors but operate on 2D frames or 2D-derived latents, leaving the geometry needed for contact-rich manipulation implicit. Paper and code are linked on GitHub.
More from Embodied
- LimX Dynamics TRON 2 humanoid shows formwork assembly and rebar tying for construction — chris_j_paxton · 2026-09-02
- Dev jokes about Physical AI: at least software only gets your drive reformatted — burhop · 2026-09-02
- Free 49,000-Word ROS 2 Hands-On Book Teaches SLAM to Autonomous Driving in Gazebo — 4310sy · 2026-09-02
- Humanoid robot attacks customer in Russian store after being pushed, staff wrestle it down — Polymarket · 2026-09-02
- WSJ: Evan Spiegel bets big on Snap's smart glasses while execs doubt internally — KateClarkTweets · 2026-09-02
- Snap to launch $2,195 smart glasses as Evan Spiegel reportedly invests big — Polymarket · 2026-09-02