FIVE-VLA Runs Driving With Just 640M Parameters, 7.5x More Efficient Than SimLingo
abursuc · x · 2026-09-17
Presenting at #ssad2026, the author showcased FIVE-VLA, a vision-language-action model with only 640M total parameters that performs surprisingly well on driving tasks.
The efficiency comes from deliberate design choices: high image resolution, a very small vision encoder and LLM, a memory encoder/decoder, and fewer tokens — resulting in 7.5x better efficiency than SimLingo on T4. A case that driving VLAs don't need to be huge; architecture trade-offs matter more.
Related event: SSAD Talk Details Building Compact Driving VLA Models(3 posts)→
More from Embodied
- GPT-Policy: In-Context Robot Learning with VLM Agents, No Gradient Updates — Dongzhou Cheng · 2026-09-17
- World Labs' Atlas Scans by Generative Guessing; NeRF Creator Admits Productization Is Hard — cen6wkf · 2026-09-17
- CXMT's LPDDR5X lands in flagship phone as Nubia ships $885 Doubao AI handset — pstAsiatech · 2026-09-17
- Key open challenges for VLAs: language, evaluation, deployment, causal reasoning — abursuc · 2026-09-17
- LADA: latent actions imitate language from few observation-language pairs — abursuc · 2026-09-17
- A Primer on Latent Action Models From the #ssad2026 Talks — abursuc · 2026-09-17