Frontier AI Agents Can Do Coding, But Fail at Open-Ended AI Research

Peter Kirgis · hf · 2026-07-30

Can current AI agents conduct open-ended AI research? This paper introduces a third evaluation method called "Shadow Evaluations" to find out.

Researchers tasked frontier AI agents with the central open-ended questions of two high-quality unpublished papers, giving them six days and thousands of dollars in compute. The agents successfully completed all engineering tasks without human help, but failed to make substantial progress on answering the core research questions, resulting in unambiguous rejections by the original authors.

The study identifies five recurring failure modes in agents during research:

The findings provide early evidence that while today's agents can handle the engineering of AI research, they still struggle with critical parts of the research lifecycle.

Original post →

More from coding & agent

coding & agent channel →