AI Agents Fail at Open-Ended Research: Authors Reject Papers After 6-Day Test
sayashk · x · 2026-07-31
Current evaluations of AI agents often focus on narrow, verifiable tasks, whereas real scientific research is open-ended. Researchers conducted a "shadow evaluation" where AI agents were given research questions from two unpublished papers, six days, and thousands of dollars in API credits and compute.
While agents proved fluent in engineering tasks, they struggled with core research skills like formulating hypotheses, determining appropriate evidence, and recognizing failing approaches. The original authors reviewed the AI-generated papers and unambiguously rejected them.
Related event: AI Agents Fail Open-Ended Research Tasks, Rejected by Original Authors(8 posts)→
More from coding & agent
- terminal-browser: Run and Control Browsers Directly Inside Your Terminal — ycombinator · 2026-07-31
- Open-Source Tmux Session Manager Registers as MCP Server, Lets Agents Name Their Own Sessions — khalon23 · 2026-07-31
- AdaMAST: Boosting AI Agent Performance with Failure Taxonomies — abeirami · 2026-07-31
- Circle and XPRIZE Launch $50K Agentic Economy Prize for Autonomous AI Payments — PeterDiamandis · 2026-07-31
- Gemini API Tip: Managed Agents Include a 7-Day Code Sandbox — tristanbob · 2026-07-31
- Managing 200+ MCP Tools: How to Avoid Agent Overload? — Calm-Republic9370 · 2026-07-31