T-RO Survey Disentangles VLA, Diffusion Policy and LLM Planner Axes in Robot Manipulation

jiqizhixin · x · 2026-09-13

A survey from Xi'an Jiaotong University, HKUST(GZ), and Peking University, accepted to IEEE Transactions on Robotics, reorganizes foundation-model-era robot manipulation from planning and learning perspectives. It argues that VLA, Diffusion Policy, Imitation Learning, and LLM Planner are compared side by side but live on different axes: VLA is about how multimodal inputs enter the action model, Diffusion Policy about how actions are generated, Imitation Learning about learning from demonstrations, and LLM Planner about high-level task planning. The paper disentangles four routinely conflated dimensions: system hierarchy, model architecture, action generation method, and learning strategy.

Original post →

More from Embodied

Embodied channel →