FULL STORY

AI Autonomous Research Test: $3000 Spent, All Papers Rejected

A recent experiment tested AI agents' ability to conduct open-ended research. Despite strong execution, the AI-generated papers were completely rejected by original authors due to a lack of metacognition.

2026-07-30 ~ 2026-08-02 · 3 episodes · 16 posts

Episode 1 · Shadow Evaluations: AI Agents' Open-Ended Research Rejected by Original Authors (2026-07-30, 12 posts)

A recent preprint introduces a testing framework called 'Shadow Evaluations' to examine whether frontier AI agents can handle open-ended scientific research. In the experiment, agents took on core open problems from two unpublished NeurIPS papers, with 6 days and thousands of dollars in API and compute resources for autonomous exploration. Ultimately, the original paper authors reviewed the AI-generated research papers and rejected them all. The study highlights clear bottlenecks for agents in open-ended research, providing an important counterweight to overly optimistic narratives of 'automated AI science'.

Confirmed

  • The team proposed the 'Shadow Evaluations' framework, where AI agents autonomously explore core open problems from two unpublished NeurIPS submissions.
  • Agents were given 6 days and thousands of dollars in compute and API credits, running hundreds of experiments.
  • Original authors rigorously reviewed the AI-generated papers and rejected all of them.
  • Agents performed reasonably on engineering/coding tasks with easily verifiable results, but struggled in open-ended research requiring hypothesis generation and evidence evaluation.
  • The study identified five recurring failure modes of agents in open-ended tasks.

Unconfirmed

  • The authors note that the negative results are preliminary, with limitations in sample size and potential for scaffolding improvements; agents' research potential remains to be further evaluated.

Why it matters

  • Current evaluations of AI agents often focus on narrow, easily verifiable tasks, whereas real scientific research is open-ended with unclear goals. This study directly addresses this gap, revealing significant bottlenecks in deep innovation tasks and providing a valuable benchmark framework for assessing agents' true research capabilities.

Episode 2 · AI Agents Lack Metacognition, Act Like Students Doing Homework (2026-07-30, 2 posts)

Experts point out that while current AI models excel at execution, they lack metacognition, acting like unreflective students who blindly chase intermediate goals without considering the overall purpose.

Episode 3 · AI Agent Autonomous Research Fails: Papers Rejected (2026-08-01, 2 posts)

An experiment testing a frontier AI agent's ability to conduct open-ended research gave it 6 days, a $3,000 budget, and sandbox access. Although the agent successfully produced two papers, both were rejected by human reviewers, highlighting that AI lacks the necessary judgment despite strong execution capabilities.