Pure VLAs may not need long-horizon planning if VLMs can cover it
m_wulfmeier · x · 2026-08-04
The post argues that pure vision-language-action models may not need to be strong at long-horizon planning, because vision-language models can cover that gap and are likely to improve more easily.
It cites OpenAI’s Rubik’s Cube work as an example: the speaker says the field already has a strong solution for the long-horizon problem, so making VLAs better at planning may be less important than improving VLMs and combining them with other components.
More from Embodied
- Concept mockup imagines a self-driving minivan with fold-out beds — LukeW · 2026-08-04
- ARPL makes llama.cpp adapt to ARM ISA and core topology at runtime — OpeningTough145 · 2026-08-04
- Multimodal spatial intelligence and physical AI pitched as the next frontier — richdotca · 2026-08-04
- Agility Robotics’ early Digit research robots helped seed China’s humanoid boom — chris_j_paxton · 2026-08-04
- Super Mario 64 demo on Vision Pro turns the screen into a portal — pvncher · 2026-08-04
- Walden Robotics says humanoids should amplify craftspeople, not replace them — adnothing · 2026-08-04