AIDE² paper: AI research agent self-improves for 8 days, beats 2-year hand-tuned harness
ptkbhv · x · 2026-09-24
Researchers released the arXiv paper on AIDE², a recursive self-improvement (RSI) system where AI research agents improve their own research efficiency.
- Core experiment: autoresearching the autoresearch agent itself for eight days
- Result: the self-improved version beats the harness the team hand-tuned for two years on held-out benchmarks, which the authors call the first experimental evidence of RSI
- New in this release: results on transfer across models and comparisons with more AI research agents
A notable attempt at turning RSI from theory into measured practice, though generality and reproducibility remain open questions.
More from coding & agent
- Running a brand with only AI agents: the Notch experiment applied to nail polish ads — azed_ai · 2026-09-24
- "The replacement you trained just became your boss" — Codex gets StackOverflow plugin — cto_junior · 2026-09-24
- Splitting sandbox base layers with Nix: fewer images, more auditable agent environments — sloppenheimer · 2026-09-24
- Devin adds native Teams support and first-party Microsoft 365 integration — DevinAI · 2026-09-24
- Dev on AI sandbox tooling: audit chronus at syscall/eBPF layer, hide extraneous tools — sloppenheimer · 2026-09-24
- Notion Engineer: Subsidized Tokens Mean You Should Switch Agents Freely Without Losing Context — nbaschez · 2026-09-24