AI Agents Fail at Open-Ended Research: Authors Reject Papers After 6-Day Test

sayashk · x · 2026-07-31

Current evaluations of AI agents often focus on narrow, verifiable tasks, whereas real scientific research is open-ended. Researchers conducted a "shadow evaluation" where AI agents were given research questions from two unpublished papers, six days, and thousands of dollars in API credits and compute.

While agents proved fluent in engineering tasks, they struggled with core research skills like formulating hypotheses, determining appropriate evidence, and recognizing failing approaches. The original authors reviewed the AI-generated papers and unambiguously rejected them.

Related event: AI Agents Fail Open-Ended Research Tasks, Rejected by Original Authors(8 posts)→

Original post →

More from coding & agent

coding & agent channel →