EmbodiedSWE: Frontier Coding Agents Crack Long-Horizon Dexterous Robotics Tasks in Sim

PeterHndrsn · x · 2026-09-25

EmbodiedSWE explores using frontier coding agents to control robots in simulation. In an IKEA table assembly demo, Fable 5.1 and Opus 5 fail while Opus 5.5 and GPT-6 Astra both succeed—Opus 5.5 finishes faster, GPT-6 Astra takes a cleaner, safer trajectory—with success rates climbing sharply over just months. The project has four parts: EmbodiedSWE-Bench, an agent-native benchmark of long-horizon dexterous everyday tasks built on Isaac Lab; evaluations of frontier coding agents on task performance, completion time, and inference cost; EmbodiedSWE-Gen, which diversifies one verified agent solution into large-scale trajectory data for training general robot policies; and RL-based agent improvement using verified outcomes, plus automatic new-task generation. Code, benchmark, and paper are open on GitHub.

Related event: EmbodiedSWE: Coding Agents Tackle Long-Horizon Robotic Tasks(4 posts)→

Original post →

More from Embodied

Embodied channel →