First AI4AI Survey Maps Why AI Can't Yet Reliably Improve AI: The Composition Gap
新智元 · wechat · 2026-09-11
The first comprehensive AI-for-AI (AI4AI) survey reviews hundreds of studies spanning long-horizon agents, automated AI R&D, and recursive self-improvement. Key findings: (1) a "composition gap"—strong per-step performance doesn't guarantee reliable end-to-end research, failures concentrate at handoffs; (2) AI executes but humans still control goals and acceptance criteria, with improvements needed on both model and harness sides; (3) claimed gains must be validated along four non-substitutable axes: measured gain, retention, human comparison, and held-out transfer. Iterated self-improvement can regress via constraint drift or metric gaming. Verdict: AI can help improve AI under explicit conditions, but durable, transferable, accumulating improvements remain unproven.
More from coding & agent
- User stumbles on mystery browser-use RL training environment, sparking agent training leak jokes — xeophon · 2026-09-11
- GPT-6 Astra + Blender MCP ships a browser-playable 3D game via WASM and WebGPU, no ThreeJS — chongdashu · 2026-09-11
- Team killed an AI agent on day one: integration quality beats model benchmarks — arthaudm · 2026-09-11
- Replacing a 3-hour Monday catchup with a weekly AI agent digest — Friendly_Parsnip_472 · 2026-09-11
- Running Qwen3.8 locally on a 128GB laptop for agentic coding: thinking tokens, not tok/s, set the wall clock — deepu105 · 2026-09-11
- Blacksmith ships free default container caching, CI jobs 25-89% faster — DanielLockyer · 2026-09-11