FIVE-VLA Runs Driving With Just 640M Parameters, 7.5x More Efficient Than SimLingo

abursuc · x · 2026-09-17

Presenting at #ssad2026, the author showcased FIVE-VLA, a vision-language-action model with only 640M total parameters that performs surprisingly well on driving tasks.

The efficiency comes from deliberate design choices: high image resolution, a very small vision encoder and LLM, a memory encoder/decoder, and fewer tokens — resulting in 7.5x better efficiency than SimLingo on T4. A case that driving VLAs don't need to be huge; architecture trade-offs matter more.

Related event: SSAD Talk Details Building Compact Driving VLA Models(3 posts)→

Original post →

More from Embodied

Embodied channel →