Qwen-RobotManip: alignment before scale for robotic manipulation foundation models

rsasaki0109 · x · 2026-09-03

The paper introduces Qwen-RobotManip, a generalizable vision-language-action foundation model built on Qwen-VL / Qwen3.5-4B. It pairs a vision-language backbone with a flow-matching Diffusion Transformer action expert for continuous action generation while preserving perception and language grounding.

The central principle is alignment before scale: robot manipulation data is inherently heterogeneous across embodiments, action spaces, cameras, coordinate frames and task distributions. The work proposes a unified alignment framework across representation, motion and behavior so multi-source training becomes coherent, trained using only open-source and egocentric datasets.

Original post →

More from Embodied

Embodied channel →