Frontier AI Agents Can Do Coding, But Fail at Open-Ended AI Research
Peter Kirgis · hf · 2026-07-30
Can current AI agents conduct open-ended AI research? This paper introduces a third evaluation method called "Shadow Evaluations" to find out.
Researchers tasked frontier AI agents with the central open-ended questions of two high-quality unpublished papers, giving them six days and thousands of dollars in compute. The agents successfully completed all engineering tasks without human help, but failed to make substantial progress on answering the core research questions, resulting in unambiguous rejections by the original authors.
The study identifies five recurring failure modes in agents during research:
- Poor judgment about the bar for publishable research
- Uncreative responses to shortcomings in research design
- Ineffective backtracking from dead ends
- Poor resource awareness
- Instruction drift
The findings provide early evidence that while today's agents can handle the engineering of AI research, they still struggle with critical parts of the research lifecycle.
More from coding & agent
- openwiki Integrates LangSmith: Auto-Generating Agent Docs from Runtime Traces — hwchase17 · 2026-07-30
- Dev Zero-Codes a 3D Cyberpunk Web Environment Using Only Opus Prompts — Dull_Film_2675 · 2026-07-30
- Adding Drama to AI Coding Assistants: Turning Tool Calls into Dial-up Sounds — enjalot · 2026-07-30
- Langostino: Open-Source Autonomous Drone with ROS2 and RL — tom_doerr · 2026-07-30
- Open-Sourcing ganfs: Automating Feature Selection with GANs — One_Crow_4710 · 2026-07-30
- Developer Insight: Building Multiplayer AI is Ultimately Permissions Hell — edgarpavlovsky · 2026-07-30