TurboVLA: 0.2B Parameter Model Achieves 32Hz Robotic Inference on RTX 4090

_akhaliq · x · 2026-07-31

Current Vision-Language-Action (VLA) models typically adopt an LLM-centric V→L→A pathway, incurring massive computation and memory overhead at every inference step.

The paper introduces TurboVLA, which reformulates this pathway into a direct V+L→A mapping. Instead of using an LLM as the central interface between perception and action, TurboVLA encodes visual and linguistic instructions independently and predicts continuous action chunks via lightweight bidirectional interaction.

On the LIBERO benchmark, TurboVLA achieves a 97.7% average success rate with only 0.2B parameters. It runs on a consumer-grade RTX 4090 with a 31.2 ms inference latency and less than 1 GB VRAM, matching or outperforming significantly larger VLA policies.

Related event: TurboVLA Enables 32Hz Robot Control on Consumer GPUs(2 posts)→

Original post →

More from Embodied

Embodied channel →