SSAD 2026: VLAs fail on new cameras, CoT doesn't reflect real causes

abursuc · x · 2026-09-18

More pessimistic results from SSAD 2026: vision-language-action models do not generalize to new cameras, and their chain-of-thought explanations do not reflect the internal cause of the action. Together with earlier claims that diffusion and public datasets fail to generalize, it points to reliability gaps in current embodied-AI approaches.

Related event: SSAD 2026: Researchers Sound Alarm on VLA and Diffusion Model Generalization(3 posts)→

Original post →

More from Embodied

Embodied channel →