T-RO Survey Disentangles VLA, Diffusion Policy and LLM Planner Axes in Robot Manipulation
jiqizhixin · x · 2026-09-13
A survey from Xi'an Jiaotong University, HKUST(GZ), and Peking University, accepted to IEEE Transactions on Robotics, reorganizes foundation-model-era robot manipulation from planning and learning perspectives. It argues that VLA, Diffusion Policy, Imitation Learning, and LLM Planner are compared side by side but live on different axes: VLA is about how multimodal inputs enter the action model, Diffusion Policy about how actions are generated, Imitation Learning about learning from demonstrations, and LLM Planner about high-level task planning. The paper disentangles four routinely conflated dimensions: system hierarchy, model architecture, action generation method, and learning strategy.
More from Embodied
- Robotics race is about teleoperation economics: US pays $150K per operator, China a third — paigeinsf · 2026-09-13
- Tesla robotaxi run-rate hits 15x in 5 days, closing in on Waymo — skorusARK · 2026-09-13
- Designer uses GPT-6 Astra with print-and-fit feedback loop to make a perfect 3D-printed enclosure — OpenAIDevs · 2026-09-13
- Someone is building a hospital for robots — _Stocko_ · 2026-09-13
- Speed spots gold Tesla robotaxi cruising Texas with no steering wheel or pedals — yunta_tsai · 2026-09-13
- Hackathon agent layer acts as an on-site engineer living inside AMR robots — ViewAdditional1739 · 2026-09-13