AI Agents Fail Open-Ended Research: Original Authors Reject All Outputs
A recent preprint introduces a "shadow evaluation" framework to test whether frontier AI agents can handle open-ended AI R&D. Agents were given 6 days, thousands of dollars in API credits, and compute to tackle core open questions from two high-quality unpublished papers. Ultimately, the original authors reviewed and explicitly rejected the AI-generated papers.
Confirmed
- Researchers proposed and utilized the "shadow evaluation" framework, tasking AI agents with autonomously exploring research questions from two unpublished NeurIPS 2026 papers.
- Agents were provided with 6 days and thousands of dollars in compute and API credits.
- The original authors peer-reviewed the AI-generated papers and rejected them outright.
- Agents performed adequately on engineering and coding tasks with easily verifiable outcomes, but struggled with open-ended research.
- The study identified five recurring failure modes for agents in open-ended research.
Unconfirmed
- The authors noted that these negative results are preliminary, with subsequent limitations in sample size and room for scaffolding improvements.
Why it matters
- This research directly addresses a core pain point in current AI development: large language models still face massive bottlenecks when handling open-ended scientific tasks requiring deep innovation and ambiguous goals. It serves as a reality check for the overly optimistic "AI automated research" narrative and provides a highly valuable benchmark framework for evaluating the true research capabilities of AI agents.
2026-07-30 ~ 2026-07-31 · 7 related posts
Primary sources
- [source] Frontier AI Agents Can Do Coding, But Fail at Open-Ended AI Research — Peter Kirgis · 2026-07-30
- [source] AI Agents Fail Open-Ended Research: Papers Rejected by Original Authors — sayashk · 2026-07-31
- [source] Agents Struggle with Open-Ended AI Research: Preprint Identifies 5 Failure Modes — random_walker · 2026-07-31
- Frontier AI Agents Fail Open-Ended Research, Papers Rejected by Authors — sayashk · 2026-07-31