USTC Study: AI Agents Achieve Only 3.3% Zero-Intervention Execution in Robot Labs

新智元 · wechat · 2026-08-06

A new study from USTC evaluates the ability of AI agents to conduct end-to-end scientific experiments in the physical world. Researchers built a robotic catalysis lab with 45 modular automated workstations, encapsulating equipment capabilities into machine-readable skills for AI to invoke.

The team stress-tested 48 configurations, combining 6 agent frameworks and 9 LLMs across 4,608 experimental trials. Results show a significant gap before AI can "take over" the lab: only 3.3% (151 workflows) were executable without human intervention. The best-performing combination of Claude Code and Claude Opus achieved an execution rate of just 28.1%.

Furthermore, while agents could adjust local parameters based on feedback, they failed to redesign research strategies or correct critical omissions in long-horizon tasks. The study concludes that fluent plan generation does not automatically translate to reliable physical execution, identifying long-horizon planning as the core bottleneck for AI research agents.

Original post →

More from Embodied

Embodied channel →