AIDE² Outperforms Hand-Tuned Version
PMinervini · x · 2026-07-15
The author highlights an experiment with AIDE²: after a research team spent two years manually tuning an autoresearch agent, a new experiment letting "agents research agents" autonomously surpassed the manual version on a held-out benchmark in just 8 days.
The key takeaway is that during 100 steps of unsupervised search, AIDE² explored many traditional approaches—like island GA, MCTS backup, and restart policies—but most were rejected. Ultimately, it won by combining simple mechanisms the team hadn't anticipated. The author views this as an early empirical signal of Recursive Self-Improvement (RSI) at Level 1.
Related event: AIDE² Self-Improvement Run Beats 2 Years of Manual Tuning(12 posts)→
More from coding & agent
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- Anthropic researcher: 99% of engineers now run swarms of 300+ self-improving agents — AlishaOutridge · 2026-09-11
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11