Princeton study: LLM coding agents beat expert-built robot planners at $20 of compute

ziv_ravid · x · 2026-10-09

A paper from Tom Silver's group at Princeton gave general LLM-based coding agents (like Claude Code and Codex) a robot task-and-motion planning problem, a simulator, and a $20 compute budget, asking whether they could out-program expert-built systems. They could — with clearly higher success rates and much faster runtimes.

Blogger Eyal Weiss then reviewed all 112 agent-written programs looking for novel algorithms. He found none. Instead, the agents did very good engineering: selecting decades-old established methods and tailoring them precisely to the task, patching in missing data where needed.

The post also explains the underlying problem: robots must jointly decide what to do (which objects to move, whether to use tools) and how to move (exact arm paths), and the two are entangled. The classic approach requires experts to hand-write action vocabularies per problem class — slow to build and slow to search.

Original post →

More from coding & agent

coding & agent channel →