Flex-π: a 6B world-action model beats π0.5 by up to 6x on real bimanual tasks

chris_j_paxton · x · 2026-09-23

Researchers from the University of Washington and Allen AI released Flex-π, a 6B-parameter world-action model that jointly denoises RGB, 3D geometry, object-centric DINO semantics and actions in a shared latent space. Per-stream dropout yields one checkpoint that runs on any subset of streams, letting developers pick a speed-accuracy operating point at deployment. On a real bimanual YAM workcell — contact-rich, sub-millimeter and long-horizon tasks including gripper self-repair — Flex-π averages 83% task completion in-distribution, beating the strongest baselines by up to 2–6x while running faster than π0.5. Paper and code are public; a RoboPapers podcast episode is coming soon.

Original post →

More from Embodied

Embodied channel →