Flex-π: 6B world action model predicts 3D pointmaps and DINO features, not just color video

chris_j_paxton · x · 2026-09-23

Flex-π is a 6-billion-parameter world action model for robots built by Ge Yan, Jesse Zhang and team. Unlike world models that only reconstruct color video, it predicts 3D pointmaps and DINO features alongside RGB, since geometry and object semantics matter more to a moving robot than color. The result is a policy that is far more demonstration-efficient, generalizes well, and handles complex, precise long-horizon tasks.

Related event: Flex-π: 6B World Action Model Predicts 3D Geometry to Boost Data Efficiency(3 posts)→

Original post →

More from Embodied

Embodied channel →