AI Agents Fail Open-Ended Research: Authors Reject Papers Generated with 6 Days & Thousands in Compute

sayashk · x · 2026-07-31

Sayash Kapoor and colleagues evaluated whether AI agents can conduct open-ended AI research, introducing a method called CRUXes (shadow evaluations). Agents were given research questions from unpublished papers, six days, and thousands of dollars in compute. The original authors then reviewed the AI-generated outputs.

The authors unambiguously rejected the agents' papers. While fluent in engineering tasks, the agents struggled with open-ended research requiring hypothesis generation, evidence selection, and recognizing failing approaches. However, researchers note that expanding these evaluations to different types of papers might soon reveal where models eventually succeed.

Related event: AI Agents Fail Open-Ended Research, Rejected by Original Authors(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →