AI Agents Fail Open-Ended Research: Authors Reject Papers After 6-Day Test

sayashk · x · 2026-07-31

Researchers conducted a "shadow evaluation" to test whether AI agents can handle open-ended AI research. Agents were given research questions from two unpublished papers, six days, and thousands of dollars in API credits and compute.

However, the original authors unambiguously rejected the AI-generated outputs after reviewing them. While agents were fluent in most engineering tasks, they lacked the persistence, ability to identify literature gaps, and keen awareness of the original research question required for true research. This also raises concerns: the point where agents discover new findings may also be the point where we can no longer catch their mistakes.

Related event: AI Agents Fail Open-Ended Research Tasks, Rejected by Original Authors(8 posts)→

Original post →

More from coding & agent

coding & agent channel →