Anthropic Tests Claude on Robotics: Direct Control Fails, LLM Supervision Hurts Familiar Tasks
DJiafei · x · 2026-07-30
Anthropic published a study on how LLMs like Claude perform on robotics tasks. They tested models controlling various embodiments—from classic control toys and a simulated quadruped to a robotic arm and a real Unitree Go2—using methods ranging from direct motor torque commands to writing controller code and providing high-level steering to pretrained policies.
Key Findings:
- Model capability depends heavily on how they are connected to the robot.
- Models mostly fail when they must drive the joints themselves.
- They can complete real navigation and manipulation tasks when supervising a pretrained controller or using simple orientation tools.
Interestingly, the study noted a counterintuitive finding highlighted by external testers: LLM supervisors can actually hurt performance on familiar tasks where a pretrained VLA model beats every LLM+VLA combo. However, for novel tasks the VLA can't solve alone, the best models provide a net uplift.
More from Embodied
- Anbernic RG Rotate meets Rabbit R1: a retro AI device mashup — SimonBalmain · 2026-07-30
- Testing Codex Agent to Autonomously Deconstruct Hardware Design and Navigate Supply Chains — mattfreed · 2026-07-30
- OpenAI President Confirms 'Family of Devices' for AI Chatbots in Development — The Verge AI · 2026-07-30
- Sneak Peek at Axol Mobile Robot: Holonomic Movement & Telescoping Lift — chris_j_paxton · 2026-07-30
- MicroFactory Launches Token to Fund Electronics Assembly Robots — ihorbeaver · 2026-07-30
- $8 Microcontroller Powers Zero-Delay Auto-Aim, Maker Wants Lasers on Lawnmower — LinusEkenstam · 2026-07-30