AI Agents Fail Open-Ended Research: Papers Rejected by Original Authors

sayashk · x · 2026-07-31

A research team explored whether AI agents can conduct open-ended AI research. Agents were given 6 days, thousands of dollars in API credits, and compute to tackle research questions from two unpublished papers. The original authors then reviewed the AI-generated papers.

The authors unambiguously rejected the agents' outputs in this "shadow evaluation."

While agents excelled at engineering tasks like literature reviews, debugging GPU environments, running hundreds of experiments, and formatting LaTeX, they fell significantly short of human researchers in core open-ended research abilities: formulating hypotheses, deciding appropriate evidence, and recognizing failing approaches.

Related event: AI Agents Fail Open-Ended Research, Rejected by Original Authors(6 posts)→

Original post →

More from coding & agent

coding & agent channel →