Flex-π: 6B world action model predicts 3D pointmaps and DINO features, not just color video
chris_j_paxton · x · 2026-09-23
Flex-π is a 6-billion-parameter world action model for robots built by Ge Yan, Jesse Zhang and team. Unlike world models that only reconstruct color video, it predicts 3D pointmaps and DINO features alongside RGB, since geometry and object semantics matter more to a moving robot than color. The result is a policy that is far more demonstration-efficient, generalizes well, and handles complex, precise long-horizon tasks.
Related event: Flex-π: 6B World Action Model Predicts 3D Geometry to Boost Data Efficiency(3 posts)→
More from Embodied
- Sim2Real robot demo makes the rounds on X — yacineMTB · 2026-09-23
- Hiro's Origin two-arm robot is being groomed to assemble its own successor's hardware — chris_j_paxton · 2026-09-23
- Three Lanes Blocked by Waymo Vehicles: A New Traffic Headache — Stup_ape_1_banana · 2026-09-23
- RealSense shows off new D585 Pro 3D stereo camera at ROSCon — chrismatthieu · 2026-09-23
- Hugging Face ships @huggingface/lerobot JS package to read robot datasets in-browser — mishig25 · 2026-09-23
- Dual Covariance 3DGS SLAM Decouples Rendering from Registration, Tracks at 60 FPS — zhenjun_zhao · 2026-09-23