AI Agents Fail at Open-Ended Research: Automation Remains Out of Reach
sudoraohacker · x · 2026-08-01
Countering optimistic visions of Recursive Self-Improvement (RSI) and automated research, researchers highlight severe limitations in current AI agents based on a recent paper and 'shadow evaluations'.
Core Experiment & Findings:
- Agents were given 6 days and thousands of dollars in API credits to replicate two unpublished papers. The original authors unambiguously rejected the AI-generated outputs.
- While fluent at engineering tasks, agents struggled significantly with open-ended research requiring high-level judgment and tacit knowledge.
Why Cutting-Edge Research is Hard to Automate:
- High Iterative Complexity: Producing a top-tier paper requires numerous iteration cycles across ideation, design, execution, writing, and review—far more complex than typical white-collar labor.
- Unpredictable Failure Modes: Delegating these processes to AI introduces dozens of emergent errors. Human experts use tacit knowledge to refine work, whereas AI models are just as likely to introduce new errors during iteration.
- Underestimation of Difficulty: Those conjecturing about the automation of scientific research severely underestimate the actual threshold for producing novel, validated scientific work.
More from coding & agent
- Agent Arena Pareto frontier: Claude and Kimi lead in cost-performance efficiency — arena · 2026-08-25
- LeanHEBO reimplements Huawei's algorithm 3x faster — hbouammar · 2026-08-25
- AI drastically reduces build time for Home Assistant configurations — HaktanSuren · 2026-08-25
- Agent runs autonomously for 24 days: System control beats pure model power — nodo48 · 2026-08-25
- Bananastand: CLI Tool to Check Real-time Value of RAM and Storage — dbreunig · 2026-08-25
- One Prompt Moves Your Coding Agent Session Across Claude, Codex, Pi and More — xhluca · 2026-08-25