AI Agents as Researchers: 6 Days, $1000, Engineering Pass but Research Fail
RexDouglass · x · 2026-08-01
A new empirical paper investigates whether AI agents can autonomously conduct open-ended scientific research. The researchers introduced a 'shadow evaluation' method: frontier AI agents were given the core open-ended questions of two high-quality unpublished NeurIPS 2026 papers, along with six days and thousands of dollars in compute.
The results show that while agents completed all engineering implementation without human help, they made no substantial progress toward answering the core research questions, leading to unambiguous rejections by the original authors.
The paper identifies five recurring failure modes: poor judgment regarding the bar for publishable research, lack of creativity, over-reliance on known methods, inability to debug complex issues effectively, and a lack of deep insight into experimental results.
More from AGI Musings
- The Jetsons Foreshadowed Home Humanoid Liability 64 Years Ago, Still Unresolved — lukas_m_ziegler · 2026-08-01
- Ford Rehires 300 Engineers After AI Fails to Replace Them — DavidLinthicum · 2026-08-01
- How Can 100 IQ Humans Control a Billion-IQ Superintelligence? — ZeroStateReflex · 2026-08-01
- Viewpoint: AI Accelerates Finding New Problems, Human Labor Remains Essential — robleclerc · 2026-08-01
- AI Pharma Bottleneck is Data, Not Models: Industry Shift Expected — MatthewMcAteer0 · 2026-08-01
- DeepSeek makes hash tables think and agentic — bronzeagepapi · 2026-08-01