SSAD 2026: VLAs fail on new cameras, CoT doesn't reflect real causes
abursuc · x · 2026-09-18
More pessimistic results from SSAD 2026: vision-language-action models do not generalize to new cameras, and their chain-of-thought explanations do not reflect the internal cause of the action. Together with earlier claims that diffusion and public datasets fail to generalize, it points to reliability gaps in current embodied-AI approaches.
More from Embodied
- Summer school dissects SOTA self-driving stacks one by one, showing they don't generalize — ftm_guney · 2026-09-18
- Robot planning heads to closed-loop settings with neural rendering for realistic environments — abursuc · 2026-09-18
- Holgercke Caesar on autonomous driving: scaling laws, end-to-end learning, and the gap — abursuc · 2026-09-18
- Hugging Face co-founder Thom Wolf shows off iterating robot head prototype — Thom_Wolf · 2026-09-18
- Open-source desk robot Sudo launches: self-hosted AI, $0 subscription, 25 founder units — wateriscoding · 2026-09-18
- LAIA dataset: 15 hours of CARLA driving with human gaze labels for explainable end-to-end AV research — abursuc · 2026-09-18