SSAD talk breaks down building a driving VLA in 4 steps: model, data, action head, finetuning

abursuc · x · 2026-09-17

At the SSAD 2026 workshop, a talk walked through building a VLA (vision-language-action) model for driving in four steps: model selection, data curation, action-head design, and multi-stage finetuning.

It opened with a brief history of the autonomous driving stack and where the field is headed next.

Related event: SSAD Talk Details Building Compact Driving VLA Models(3 posts)→

Original post →

More from Embodied

Embodied channel →