Princeton study: LLM coding agents beat expert-built robot planners at $20 of compute
ziv_ravid · x · 2026-10-09
A paper from Tom Silver's group at Princeton gave general LLM-based coding agents (like Claude Code and Codex) a robot task-and-motion planning problem, a simulator, and a $20 compute budget, asking whether they could out-program expert-built systems. They could — with clearly higher success rates and much faster runtimes.
Blogger Eyal Weiss then reviewed all 112 agent-written programs looking for novel algorithms. He found none. Instead, the agents did very good engineering: selecting decades-old established methods and tailoring them precisely to the task, patching in missing data where needed.
The post also explains the underlying problem: robots must jointly decide what to do (which objects to move, whether to use tools) and how to move (exact arm paths), and the two are entangled. The classic approach requires experts to hand-write action vocabularies per problem class — slow to build and slow to search.
More from coding & agent
- 20 AI agent memory tools to know in 2026, sorted into five selection categories — MaryamMiradi · 2026-10-09
- scikit-learn creator on coding agents: 'They get you to do stupid things faster with more energy' — GaelVaroquaux · 2026-10-09
- Matt Pocock calls out Anthropic's plugin marketplace for stranding users on stale versions — mattpocockuk · 2026-10-09
- Alook: open-source rooms where local coding agents get handles, inboxes, and DMs — rohanpaul_ai · 2026-10-09
- Alook gives coding agents inboxes and DMs, letting Claude Code hand work to Codex — rohanpaul_ai · 2026-10-09
- TypeSafe AI's Jev decision model claims 200x faster, 400x cheaper classification in agent loops — LangChain · 2026-10-09