Flex-π: a 6B world-action model that predicts 3D geometry for data-efficient robots
chris_j_paxton · x · 2026-09-23
Chris Paxton's RoboPapers episode 106 (now with an auto-generated, ChatGPT-corrected transcript) covers Flex-π, a 6B-parameter world-action model. Key points:
- Most WAMs predict only RGB latents, though 3D geometry and object semantics matter more for robots than color.
- Flex-π exploits a "free lunch": the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps nearly losslessly, with no pointmap-specific training.
- This enables joint supervision of RGB, 3D pointmaps, and DINO object features with no new sensors, pre-training, or inference latency.
- Signals are projected into a shared latent space and denoised jointly with actions in a Mixture-of-Transformers backbone, with per-stream dropout and cross-modality forcing.
- Result: far better demonstration efficiency and generalization to complex, precise long-horizon tasks.
Related event: Flex-π: 6B World Action Model Predicts 3D Geometry to Boost Data Efficiency(3 posts)→
More from Embodied
- Robotics data startup Eidon shuts down after 2+ years, open-sources its <$30 finger tracker — vanstriendaniel · 2026-09-23
- TUM survey unifies physics-embedded robot learning with a new taxonomy — TUM-AVS · 2026-09-23
- RealSense and NVIDIA Robotics AI demos impress at ROSCon — chrismatthieu · 2026-09-23
- Joga: a full-stack humanoid soccer system with active vision and residual sim-to-real models — ericjang11 · 2026-09-23
- GLIDE has an LLM write failure-filtering guardrails, lifting robot task success from 0% to 70% — stepjamUK · 2026-09-23
- Ondas acquisitions add sensing, resilient connectivity and GNSS-denied navigation — CeoOndas · 2026-09-23