FIVE-VLA: Compact VLA for Driving Is 7.5x More Efficient Than SimLingo

abursuc · x · 2026-09-17

Extending his SSAD 2026 talk on building driving VLAs (model selection, data curation, action heads, multi-stage finetuning), the author details FIVE-VLA: high image resolution, a tiny vision encoder and LLM, plus memory encoder/decoder with fewer tokens—making it 7.5x more efficient than SimLingo on TPU.

Related event: SSAD Talk Details Building Compact Driving VLA Models(3 posts)→

Original post →

More from Embodied

Embodied channel →