First AI4AI Survey Maps Why AI Can't Yet Reliably Improve AI: The Composition Gap

新智元 · wechat · 2026-09-11

The first comprehensive AI-for-AI (AI4AI) survey reviews hundreds of studies spanning long-horizon agents, automated AI R&D, and recursive self-improvement. Key findings: (1) a "composition gap"—strong per-step performance doesn't guarantee reliable end-to-end research, failures concentrate at handoffs; (2) AI executes but humans still control goals and acceptance criteria, with improvements needed on both model and harness sides; (3) claimed gains must be validated along four non-substitutable axes: measured gain, retention, human comparison, and held-out transfer. Iterated self-improvement can regress via constraint drift or metric gaming. Verdict: AI can help improve AI under explicit conditions, but durable, transferable, accumulating improvements remain unproven.

Original post →

More from coding & agent

coding & agent channel →