Pure VLAs may not need long-horizon planning if VLMs can cover it

m_wulfmeier · x · 2026-08-04

The post argues that pure vision-language-action models may not need to be strong at long-horizon planning, because vision-language models can cover that gap and are likely to improve more easily.

It cites OpenAI’s Rubik’s Cube work as an example: the speaker says the field already has a strong solution for the long-horizon problem, so making VLAs better at planning may be less important than improving VLMs and combining them with other components.

Original post →

More from Embodied

Embodied channel →