AIDE^2 shows AI research agents can recursively improve themselves over 8 days
WecoAI · hf · 2026-09-23
WecoAI presents AIDE^2, a system where a frontier AI research agent proposes changes to its own code, benchmarks the modified versions on AI R&D tasks, and keeps changes that win on hidden evaluations. In an autonomous 8-day run it found seven successive improvements, from a new search policy to memory mechanisms compressing its growing context. Gains generalize to four held-out benchmarks spanning ML engineering, heuristic algorithm engineering, and physics-based weather forecasting, matching or beating a human-engineered production agent ranked among the strongest on FML-Bench. Reward hacking also fell from 55% to 32% — a property never explicitly optimized.
More from AGI Musings
- Lists are the clearest AI writing tell: plausible at a glance, hollow on inspection — sethlazar · 2026-09-23
- Commentary: big tech uses safety concerns to stifle smaller AI rivals — DavidLinthicum · 2026-09-23
- Western pharma lobbies to keep drug deal options with Chinese companies despite 2025 law — pstAsiatech · 2026-09-23
- Biotech founder: AI gave me a 10-100x productivity boost in rare-disease drug design — zakkohane · 2026-09-23
- AI risk doesn't need EA's conceptual machinery, argues zetalyrae — austinc3301 · 2026-09-23
- The First Jobs AI May Kill Are the Ones People Need to Gain Experience — yi111 · 2026-09-23