Frontier AI Agents Fail Open-Ended Research: $3,000 and 6 Days Yield Rejected Papers
gerardsans · x · 2026-08-02
A recent study tested whether frontier AI agents can conduct open-ended AI research. Agents were given six days, a $3,000 budget, full virtual sandboxes, and web access to tackle central research questions from two unpublished NeurIPS 2026 papers.
The results indicate that while agents completed the full execution loop with zero human intervention and handled many engineering tasks, they failed at core research objectives, and recursive self-improvement (RSI) never materialized. Original authors acting as expert reviewers rejected the outputs with low scores of 1/6 and 2/6. The study identifies five recurring failure modes, most notably poor judgment, such as misjudging task difficulty.
Related event: AI Agent Autonomous Research Fails: Papers Rejected(2 posts)→
More from coding & agent
- Agent runs autonomously for 24 days: System control beats pure model power — nodo48 · 2026-08-25
- Bananastand: CLI Tool to Check Real-time Value of RAM and Storage — dbreunig · 2026-08-25
- One Prompt Moves Your Coding Agent Session Across Claude, Codex, Pi and More — xhluca · 2026-08-25
- session-migrate: move coding agent sessions across Claude Code, Codex and more — xhluca · 2026-08-25
- Graph Engineering Guide: Building Road Maintenance AI Agents — MaryamMiradi · 2026-08-25
- Exiting Cursor Engineer Details Composer 2 Training and Kernel Design — eliebakouch · 2026-08-25