AIDE² Outperforms Hand-Tuned Version
PMinervini · x · 2026-07-15
The author highlights an experiment with AIDE²: after a research team spent two years manually tuning an autoresearch agent, a new experiment letting "agents research agents" autonomously surpassed the manual version on a held-out benchmark in just 8 days.
The key takeaway is that during 100 steps of unsupervised search, AIDE² explored many traditional approaches—like island GA, MCTS backup, and restart policies—but most were rejected. Ultimately, it won by combining simple mechanisms the team hadn't anticipated. The author views this as an early empirical signal of Recursive Self-Improvement (RSI) at Level 1.
Related event: AIDE² Self-Improvement Run Beats 2 Years of Manual Tuning(12 posts)→
More from coding & agent
- Building a Secure AI Agent Gateway: Self-Hosting OAuth for Multiple SaaS Apps — Defiant_Cod_2654 · 2026-07-22
- Rowboat launches as an open-source, local-first AI coworker with memory — ycombinator · 2026-07-22
- Scoble says AI “loops” really means long-running multi-agent workspaces — Scobleizer · 2026-07-22
- Kimi Code opens a waitlist as Moonshot rolls out its coding product — Fabulous_Bonus_8981 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- Indie Dev Asks: What's Actually Broken in Your AI Agent's Memory Today? — AcceptableTime7937 · 2026-07-22