Sentdex shows general multimodal LLMs can drive robots with zero training
Sentdex · x · 2026-09-04
Sentdex demonstrates a contrarian approach: skip VLA training entirely and let general multimodal LLMs handle high-level robot intelligence. He wired vision-capable GLM 5.3 Flash to an XGO mini wheeled quadruped's SDK — the model analyzed camera frames and autonomously called movement, arm and gripper APIs to solve tasks, with zero training or fine-tuning. DSV4F + Qwen 3.8 27B works too, and even Z AI was surprised.
He argues the hard part of robotics isn't object detection but intelligence and planning for real-world imperfection. From experience, VLAs are finicky, sim2real-fragile, and yield single-task robots after weeks of work. For quadrupeds that don't need fast IMU loops, general LLMs may already suffice — possibly sidestepping the whole VLA/world-model research direction.
Related event: Sentdex drives a quadruped robot with a general multimodal LLM(3 posts)→
More from Embodied
- Perceptron's Isaac 0.5 robot folds t-shirts, ships open weights for cross-embodiment repro — iamrobotbear · 2026-09-04
- Lab lessons from Anthropic MHS: keep fast control out of the model — Empty-Abalone-2952 · 2026-09-04
- Build or OEM? Nvidia Fabless Lesson for Humanoid Robot Makers — chris_j_paxton · 2026-09-04
- "Atlas <> Palantir" teaser surfaces, pointing to a September 10, 2026 reveal — eliano · 2026-09-04
- Ultra's hybrid robotic-arm plus semi-humanoid strategy targets rapid 3PL field deployment — chris_j_paxton · 2026-09-04
- Zeroth Robotics launches Bridge, an 88cm humanoid for developers under $4k with open SDK — chris_j_paxton · 2026-09-04