AI Agents Fail Open-Ended Research: Authors Reject Papers Generated with 6 Days & Thousands in Compute
sayashk · x · 2026-07-31
Sayash Kapoor and colleagues evaluated whether AI agents can conduct open-ended AI research, introducing a method called CRUXes (shadow evaluations). Agents were given research questions from unpublished papers, six days, and thousands of dollars in compute. The original authors then reviewed the AI-generated outputs.
The authors unambiguously rejected the agents' papers. While fluent in engineering tasks, the agents struggled with open-ended research requiring hypothesis generation, evidence selection, and recognizing failing approaches. However, researchers note that expanding these evaluations to different types of papers might soon reveal where models eventually succeed.
Related event: AI Agents Fail Open-Ended Research, Rejected by Original Authors(9 posts)→
More from AGI Musings
- Nobel Laureate Simon Johnson on AI Job Displacement and China's Over-Automation Risks — davidyin44 · 2026-07-31
- Recursive Self-Improvement Gets a Body: How Embodied AI Could Rebuild Civilization — imjustnewatai · 2026-07-31
- Nvidia's Theoretical 2035 Valuation Could Exceed Total Human History Value — ns123abc · 2026-07-31
- Reflecting on OpenAI's Rubik's Cube Demo: The Industry Lost Its Appetite for Exploratory Research — avt_im · 2026-07-31
- Zuckerberg Outlines Meta's Superintelligence Vision: Individual Empowerment Over Concentrated Control — clu_cheng · 2026-07-31
- Exploring AI Consciousness: Listen to Models Instead of Forcing Alignment — repligate · 2026-07-31