AI Scientists Flunk Real-World Lab Tests: Only 3.3% Workflows Executable

新智元 · wechat · 2026-07-31

A new study from USTC built a robotic catalysis lab with 45 modular stations to test if AI agents can conduct end-to-end scientific discovery in the physical world. The system translates lab capabilities into machine-readable skills for AI to invoke.

Researchers stress-tested 48 configurations combining 6 agent frameworks and 9 LLMs across 4,608 trials. Results show a significant gap before AI can "take over": only 3.3% (151) of generated workflows were executable without human intervention. The best-performing combo (Claude Code + Claude 3.5 Sonnet) achieved a 28.1% execution rate.

Closed-loop tests revealed that while agents can adjust local parameters based on experimental feedback, they fail to redesign research strategies or spot critical omissions at a scientific level. Long-horizon planning remains a major bottleneck. The study distinguishes three distinct AI capabilities: generating plans, creating physically executable workflows, and adjusting overall research strategies.

Original post →

More from coding & agent

coding & agent channel →