Can AI Agents Do Open-Ended Research? Authors Reject Outputs

sayashk · x · 2026-07-31

While most agent evaluations focus on narrow, verifiable tasks, real AI research is open-ended. Researchers conducted a "shadow evaluation" giving agents research questions from two unpublished papers, six days, and thousands of dollars in API credits.

Although agents were fluent at most engineering tasks, the original authors unambiguously rejected the AI-generated papers. This highlights that AI is not yet capable of independently generating valuable new ideas (like Chinchilla scaling laws or MoE), with modal predictions for this milestone pointing to 2027.

Related event: AI Agents Fail Open-Ended Research Tasks, Rejected by Original Authors(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →