FULL STORY
AI Autonomous Research Test: $3000 Spent, All Papers Rejected
A recent experiment tested AI agents' ability to conduct open-ended research. Despite strong execution, the AI-generated papers were completely rejected by original authors due to a lack of metacognition.
2026-07-30 ~ 2026-08-02 · 3 episodes · 16 posts
Episode 1 · Shadow Evaluations: AI Agents' Open-Ended Research Rejected by Original Authors (2026-07-30, 12 posts)
A recent preprint introduces a testing framework called 'Shadow Evaluations' to examine whether frontier AI agents can handle open-ended scientific research. In the experiment, agents took on core open problems from two unpublished NeurIPS papers, with 6 days and thousands of dollars in API and compute resources for autonomous exploration. Ultimately, the original paper authors reviewed the AI-generated research papers and rejected them all. The study highlights clear bottlenecks for agents in open-ended research, providing an important counterweight to overly optimistic narratives of 'automated AI science'.
Confirmed
- The team proposed the 'Shadow Evaluations' framework, where AI agents autonomously explore core open problems from two unpublished NeurIPS submissions.
- Agents were given 6 days and thousands of dollars in compute and API credits, running hundreds of experiments.
- Original authors rigorously reviewed the AI-generated papers and rejected all of them.
- Agents performed reasonably on engineering/coding tasks with easily verifiable results, but struggled in open-ended research requiring hypothesis generation and evidence evaluation.
- The study identified five recurring failure modes of agents in open-ended tasks.
Unconfirmed
- The authors note that the negative results are preliminary, with limitations in sample size and potential for scaffolding improvements; agents' research potential remains to be further evaluated.
Why it matters
- Current evaluations of AI agents often focus on narrow, easily verifiable tasks, whereas real scientific research is open-ended with unclear goals. This study directly addresses this gap, revealing significant bottlenecks in deep innovation tasks and providing a valuable benchmark framework for assessing agents' true research capabilities.
- Frontier AI Agents Can Do Coding, But Fail at Open-Ended AI Research — Peter Kirgis · 2026-07-30
- AI Agents Fail Open-Ended Research: Papers Rejected by Original Authors — sayashk · 2026-07-31
- Agents Struggle with Open-Ended AI Research: Preprint Identifies 5 Failure Modes — random_walker · 2026-07-31
- AI Agents Fail Open-Ended Research: Authors Reject Papers After 6-Day Test — sayashk · 2026-07-31
- Frontier AI Agents Fail at Open-Ended Research Despite 6 Days of Compute — sayashk · 2026-07-31
- Frontier AI Agents Fail Open-Ended Research, Papers Rejected by Authors — sayashk · 2026-07-31
- Can AI Agents Do Open-Ended Research? Authors Reject Outputs — sayashk · 2026-07-31
- AI Agents Fail at Open-Ended Research: Authors Reject Papers After 6-Day Test — sayashk · 2026-07-31
- AI Agents Fail Open-Ended Research: Authors Reject Papers Generated with 6 Days & Thousands in Compute — sayashk · 2026-07-31
- Can AI Agents Conduct Open-Ended Scientific Research? — ChenhaoTan · 2026-08-01
- AI Agents as Researchers: 6 Days, $1000, Engineering Pass but Research Fail — RexDouglass · 2026-08-01
- AI Agents Fail at Open-Ended Research: Automation Remains Out of Reach — sudoraohacker · 2026-08-01
Episode 2 · AI Agents Lack Metacognition, Act Like Students Doing Homework (2026-07-30, 2 posts)
Experts point out that while current AI models excel at execution, they lack metacognition, acting like unreflective students who blindly chase intermediate goals without considering the overall purpose.
- AI Excels at Academic Tasks But Blindly Follows Instructions Like an Unthinking Student — davidmanheim · 2026-07-30
- AI Agents Execute Tasks Well but Lack Metacognition, Leading to Goal Misalignment — davidmanheim · 2026-07-30
Episode 3 · AI Agent Autonomous Research Fails: Papers Rejected (2026-08-01, 2 posts)
An experiment testing a frontier AI agent's ability to conduct open-ended research gave it 6 days, a $3,000 budget, and sandbox access. Although the agent successfully produced two papers, both were rejected by human reviewers, highlighting that AI lacks the necessary judgment despite strong execution capabilities.
- AI Agents Wrote Two Papers in 6 Days for $3K, Both Rejected for Lack of Judgment — rohanpaul_ai · 2026-08-01
- Frontier AI Agents Fail Open-Ended Research: $3,000 and 6 Days Yield Rejected Papers — gerardsans · 2026-08-02