Sentdex shows general multimodal LLMs can drive robots with zero training

Sentdex · x · 2026-09-04

Sentdex demonstrates a contrarian approach: skip VLA training entirely and let general multimodal LLMs handle high-level robot intelligence. He wired vision-capable GLM 5.3 Flash to an XGO mini wheeled quadruped's SDK — the model analyzed camera frames and autonomously called movement, arm and gripper APIs to solve tasks, with zero training or fine-tuning. DSV4F + Qwen 3.8 27B works too, and even Z AI was surprised.

He argues the hard part of robotics isn't object detection but intelligence and planning for real-world imperfection. From experience, VLAs are finicky, sim2real-fragile, and yield single-task robots after weeks of work. For quadrupeds that don't need fast IMU loops, general LLMs may already suffice — possibly sidestepping the whole VLA/world-model research direction.

Related event: Sentdex drives a quadruped robot with a general multimodal LLM(3 posts)→

Original post →

More from Embodied

Embodied channel →