AI Agents Fail Open-Ended Research: Original Authors Reject All Outputs

A recent preprint introduces a "shadow evaluation" framework to test whether frontier AI agents can handle open-ended AI R&D. Agents were given 6 days, thousands of dollars in API credits, and compute to tackle core open questions from two high-quality unpublished papers. Ultimately, the original authors reviewed and explicitly rejected the AI-generated papers.

Confirmed

Unconfirmed

Why it matters

2026-07-30 ~ 2026-07-31 · 7 related posts

Primary sources

3 near-duplicate retellings: sayashk · sayashk · sayashk