Same model, two placements: agentic control completes 18/20 LEGO tasks vs 6/20 for code-as-policy

paigeinsf · x · 2026-09-24

SE3 Labs ran a controlled comparison of where a robot's intelligence should live, using the same OpenAI GPT6-Astra model, the same bimanual YAM station, and the same LEGO pick-and-place task on a real-world remote eval stack:

The gap wasn't grasping but recovery: code-as-policy lifted the brick in 14/20 episodes yet failed the rest without retrying; the agent needed multiple grasp attempts in 9 of its 18 completions, including one episode where it dropped the brick, reacquired it, and placed it at 419s. The tradeoff is speed: median release time 175s for agentic vs 131s for code-as-policy among successful runs, with the agent also granted up to 600s to recover vs 180s. The report includes a 0–100 progress rubric and a 100-point quality score (grasp 30 / transport 30 / placement 40), scored on-site by a human operator. Full report and all 40 episode videos are public.

Original post →

More from Embodied

Embodied channel →