Anthropic Tests Claude on Robotics: Direct Control Fails, LLM Supervision Hurts Familiar Tasks

DJiafei · x · 2026-07-30

Anthropic published a study on how LLMs like Claude perform on robotics tasks. They tested models controlling various embodiments—from classic control toys and a simulated quadruped to a robotic arm and a real Unitree Go2—using methods ranging from direct motor torque commands to writing controller code and providing high-level steering to pretrained policies.

Key Findings:

Interestingly, the study noted a counterintuitive finding highlighted by external testers: LLM supervisors can actually hurt performance on familiar tasks where a pretrained VLA model beats every LLM+VLA combo. However, for novel tasks the VLA can't solve alone, the best models provide a net uplift.

Original post →

More from Embodied

Embodied channel →