AI Agents as Researchers: 6 Days, $1000, Engineering Pass but Research Fail

RexDouglass · x · 2026-08-01

A new empirical paper investigates whether AI agents can autonomously conduct open-ended scientific research. The researchers introduced a 'shadow evaluation' method: frontier AI agents were given the core open-ended questions of two high-quality unpublished NeurIPS 2026 papers, along with six days and thousands of dollars in compute.

The results show that while agents completed all engineering implementation without human help, they made no substantial progress toward answering the core research questions, leading to unambiguous rejections by the original authors.

The paper identifies five recurring failure modes: poor judgment regarding the bar for publishable research, lack of creativity, over-reliance on known methods, inability to debug complex issues effectively, and a lack of deep insight into experimental results.

Related event: Shadow Evaluations: AI Agents Fail Open-Ended Research, Rejected by Original Authors(11 posts)→

Original post →

More from AGI Musings

AGI Musings channel →