IMLE-VLA replaces flow matching with single-step cIMLE, 3.67x faster robot actions at 55Hz
petitegeek · x · 2026-09-28
Researchers from SFU and UPenn introduced IMLE-VLA, tackling the slow, jerky inference of vision-language-action (VLA) policies.
- Key change: Existing VLAs like π0.5 use a 10-step flow-matching action head; IMLE-VLA swaps it for a single-step conditional IMLE (cIMLE) generator, keeping the same VLM backbone but requiring one forward pass instead of ten.
- Results: Inference jumps from 15Hz to 55Hz (3.67x), with up to 11x higher action throughput. It achieves the highest average success rate on the 40-task LIBERO benchmark (98.0%).
- Theory: cIMLE provably preserves multimodal action coverage, avoiding mode collapse without multi-step sampling.
- Real-world: Franka Panda experiments show 2.2–3.0x lower jerk and 3.9–6.6x lower VLA compute time, with smoother motion and much faster task completion.
Paper and code are open source; the work will be presented at IROS 2026.
More from Embodied
- AnyMo, a Setup-Agnostic IMU Human Motion Model, Accepted at NeurIPS 2026 — flosalim · 2026-09-28
- Leslie Kaelbling on Why Astra Excels at Robotics—and Why It Isn't Solved Yet — shaohua0116 · 2026-09-28
- Meta's official XR FOV Simulator used for accurate Quest 3 vs VR glasses FOV comparison — ZeroStateReflex · 2026-09-28
- Robotics researcher argues proprioception, not vision, is the real bottleneck for grasping — arankomatsuzaki · 2026-09-28
- Dev builds own Playdate visual simulator with Figma-to-dither workflow — eschadiol · 2026-09-28
- InternW0-Δ open-sources code, weights and 20K+ hours of robot data — arankomatsuzaki · 2026-09-28