AI Scientists Flunk Real-World Lab Tests: Only 3.3% Workflows Executable
新智元 · wechat · 2026-07-31
A new study from USTC built a robotic catalysis lab with 45 modular stations to test if AI agents can conduct end-to-end scientific discovery in the physical world. The system translates lab capabilities into machine-readable skills for AI to invoke.
Researchers stress-tested 48 configurations combining 6 agent frameworks and 9 LLMs across 4,608 trials. Results show a significant gap before AI can "take over": only 3.3% (151) of generated workflows were executable without human intervention. The best-performing combo (Claude Code + Claude 3.5 Sonnet) achieved a 28.1% execution rate.
Closed-loop tests revealed that while agents can adjust local parameters based on experimental feedback, they fail to redesign research strategies or spot critical omissions at a scientific level. Long-horizon planning remains a major bottleneck. The study distinguishes three distinct AI capabilities: generating plans, creating physically executable workflows, and adjusting overall research strategies.
More from coding & agent
- 10k-Star Reverse-Skill: AI Routing Pack for Pentesting in Coding Agents — zhaoxuya520 · 2026-07-31
- Benchmarking Kimi K3, GLM 5.2, and DeepSeek V4 Pro in Agent Workflows — Teknium · 2026-07-31
- Open-Source KG Extractor Runs Qwen on a Single NVIDIA L4 — JeremyCMorgan · 2026-07-31
- Anthropic Engineers Run Hundreds of Agents via Graph Engineering — blaizedsouza · 2026-07-31
- Build and Deploy Websites from Scratch Using Claude Code and MCP Connectors — TawohAwa · 2026-07-31
- Solving Mid-Task Agent Failures: Open-Source Tool for Automatic Rollbacks — Hour-Bite8746 · 2026-07-31