TurboVLA: 0.2B Parameter Model Achieves 32Hz Robotic Inference on RTX 4090
_akhaliq · x · 2026-07-31
Current Vision-Language-Action (VLA) models typically adopt an LLM-centric V→L→A pathway, incurring massive computation and memory overhead at every inference step.
The paper introduces TurboVLA, which reformulates this pathway into a direct V+L→A mapping. Instead of using an LLM as the central interface between perception and action, TurboVLA encodes visual and linguistic instructions independently and predicts continuous action chunks via lightweight bidirectional interaction.
On the LIBERO benchmark, TurboVLA achieves a 97.7% average success rate with only 0.2B parameters. It runs on a consumer-grade RTX 4090 with a 31.2 ms inference latency and less than 1 GB VRAM, matching or outperforming significantly larger VLA policies.
Related event: TurboVLA Enables 32Hz Robot Control on Consumer GPUs(2 posts)→
More from Embodied
- Google DeepMind Launches Gemini Robotics 2 for Whole-Body Control and Multi-Robot Collaboration — CyberRobooo · 2026-07-31
- DeepMind's Gemini Robotics 2 Enables Multi-Robot Collaboration — CyberRobooo · 2026-07-31
- YC-Backed WaveSight: A New Camera That Sees Through Walls — ycombinator · 2026-07-31
- AI Pendant Friend Relaunches with a Speaker, Doubles in Price — The Verge AI · 2026-07-31
- New MT3 Paradigm Enables Robots to Learn 1000 Tasks in Under 24 Hours — chris_j_paxton · 2026-07-31
- Simulated LiDAR Needs Real-World Noise for Better Robustness — Sentdex · 2026-07-31