FIVE-VLA: Compact VLA for Driving Is 7.5x More Efficient Than SimLingo
abursuc · x · 2026-09-17
Extending his SSAD 2026 talk on building driving VLAs (model selection, data curation, action heads, multi-stage finetuning), the author details FIVE-VLA: high image resolution, a tiny vision encoder and LLM, plus memory encoder/decoder with fewer tokens—making it 7.5x more efficient than SimLingo on TPU.
Related event: SSAD Talk Details Building Compact Driving VLA Models(3 posts)→
More from Embodied
- Ex-Tsinghua/BIGAI team's generalist dexterity policy beats Astra on RoboDojo — chris_j_paxton · 2026-09-17
- GPT-Policy: In-Context Robot Learning with VLM Agents, No Gradient Updates — Dongzhou Cheng · 2026-09-17
- World Labs' Atlas Scans by Generative Guessing; NeRF Creator Admits Productization Is Hard — cen6wkf · 2026-09-17
- CXMT's LPDDR5X lands in flagship phone as Nubia ships $885 Doubao AI handset — pstAsiatech · 2026-09-17
- Key open challenges for VLAs: language, evaluation, deployment, causal reasoning — abursuc · 2026-09-17
- LADA: latent actions imitate language from few observation-language pairs — abursuc · 2026-09-17