Frontier AI Agents Fail at Open-Ended Research Despite 6 Days of Compute

sayashk · x · 2026-07-31

A new study tests whether frontier AI agents can conduct open-ended AI R&D. The researchers introduced "shadow evaluations," where agents tackle the core open-ended questions of unpublished NeurIPS 2026 papers, given 6 days and thousands of dollars in compute.

Results show that while agents completed all engineering tasks autonomously, they failed to make substantial progress on the research questions, leading to unambiguous rejections by the original authors. The paper identifies five recurring failure modes, most notably poor judgment regarding the bar for publishable research. This provides early evidence against the assumption that AI agents are ready to automate explosive AI progress.

Related event: AI Agents Fail Open-Ended Research: Original Authors Reject All Outputs(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →