Key open challenges for VLAs: language, evaluation, deployment, causal reasoning
abursuc · x · 2026-09-17
At #ssad2026, abursuc outlined the key open challenges for vision-language-action (VLA) models:
- How much language is enough
- How to evaluate VLAs
- How to deploy them
- How to enforce causal reasoning
- How to finetune without crippling existing capabilities
- How to fix language collapse
He also offered a potential explanation for why LADA works: it breaks a difficult problem into two simpler ones.
More from Embodied
- GPT-Policy: In-Context Robot Learning with VLM Agents, No Gradient Updates — Dongzhou Cheng · 2026-09-17
- World Labs' Atlas Scans by Generative Guessing; NeRF Creator Admits Productization Is Hard — cen6wkf · 2026-09-17
- CXMT's LPDDR5X lands in flagship phone as Nubia ships $885 Doubao AI handset — pstAsiatech · 2026-09-17
- LADA: latent actions imitate language from few observation-language pairs — abursuc · 2026-09-17
- A Primer on Latent Action Models From the #ssad2026 Talks — abursuc · 2026-09-17
- FIVE-VLA Runs Driving With Just 640M Parameters, 7.5x More Efficient Than SimLingo — abursuc · 2026-09-17