Coding Agents Beat Hand-Engineered Planners at Generalized TAMP, 56%-95% vs 47%

FBK-NLP · hf · 2026-09-25

FBK-NLP tested whether coding agents can automate generalized task and motion planning (TAMP): given a task description and simulator access, agents interact with the environment to synthesize a program, which is then frozen and evaluated on unseen instances.

Code and full prompts are open-sourced. Coding agents are a strong baseline for generalized TAMP.

Original post →

More from coding & agent

coding & agent channel →