HUST & Huawei's TurboVLA Bypasses LLM for 32Hz Real-Time Robot Control

机器之心 · wechat · 2026-08-03

Existing Vision-Language-Action (VLA) models typically use Large Language Models (LLMs) as an intermediary between vision and action, resulting in high computational overhead and latency. To address this, Huazhong University of Science and Technology (HUST) and Huawei proposed TurboVLA, a real-time VLA model.

Core Architectural Innovation

TurboVLA bypasses the LLM interface, adopting a more direct V+L→A path:

Performance & Efficiency

This research proves that for explicit robotic tasks, action prediction doesn't always require running a full LLM. A layered system where LLM handles high-level planning and TurboVLA handles low-latency execution could be a more efficient technical route.

Original post →

More from Embodied

Embodied channel →