Frontier Agents Given 6 Days & Thousands in Compute Fail Core NeurIPS Research
billhilf · x · 2026-08-04
A new study tested whether frontier AI agents can conduct open-ended AI research. Agents were given six days and thousands of dollars in compute to solve the core research questions of two unpublished NeurIPS 2026 papers.
Results showed that while agents completed all engineering tasks without human help, they failed to make substantial progress on the actual research questions, leading to unambiguous rejections by the original authors. The paper identifies five recurring failure modes, including poor judgment on publishable standards and lack of creativity, highlighting the remaining gap in AI R&D automation.
More from coding & agent
- AI Makes Implementation Cheap, but Engineering Judgment Remains Priceless — haltakov · 2026-08-04
- Vibe-Coded Arena FPS: Dev Builds Browser 3D Game Using Multiple LLMs — rudesssolo · 2026-08-04
- Scalar Field Launches Agentic Trading Desk Connecting Research to Live Execution — ycombinator · 2026-08-04
- LangSmith LLM Gateway Enters Public Beta with BYOK and Anti-Lock-in Features — LangChain · 2026-08-04
- Vibe Coding in Virtual Reality on the Apple Vision Pro — nptacek · 2026-08-04
- Gemini Spark Update: Browser Agent Automates Complex Errands via Chrome — Rare_Bunch4348 · 2026-08-04