Frontier AI Agents Can Do Engineering But Fail at Open-Ended Research

TheTuringPost · x · 2026-08-08

A new study tested whether frontier AI agents can conduct open-ended AI research. Given six days and thousands of dollars in compute to solve central questions of two unpublished NeurIPS 2026 papers, the agents completed all engineering tasks without human help but failed to make substantial research progress, leading to outright rejection by the authors.

The paper identifies five recurring failure modes:

The conclusion offers a reality check: running the research loop and actually doing good research are still two very different things.

Related event: Study Shows Frontier AI Agents Struggle with Open-Ended Research(3 posts)→

Original post →

More from coding & agent

coding & agent channel →