Frontier Agents Given 6 Days & Thousands in Compute Fail Core NeurIPS Research

billhilf · x · 2026-08-04

A new study tested whether frontier AI agents can conduct open-ended AI research. Agents were given six days and thousands of dollars in compute to solve the core research questions of two unpublished NeurIPS 2026 papers.

Results showed that while agents completed all engineering tasks without human help, they failed to make substantial progress on the actual research questions, leading to unambiguous rejections by the original authors. The paper identifies five recurring failure modes, including poor judgment on publishable standards and lack of creativity, highlighting the remaining gap in AI R&D automation.

Original post →

More from coding & agent

coding & agent channel →