AIDE^2 shows AI research agents can recursively improve themselves over 8 days

WecoAI · hf · 2026-09-23

WecoAI presents AIDE^2, a system where a frontier AI research agent proposes changes to its own code, benchmarks the modified versions on AI R&D tasks, and keeps changes that win on hidden evaluations. In an autonomous 8-day run it found seven successive improvements, from a new search policy to memory mechanisms compressing its growing context. Gains generalize to four held-out benchmarks spanning ML engineering, heuristic algorithm engineering, and physics-based weather forecasting, matching or beating a human-engineered production agent ranked among the strongest on FML-Bench. Reward hacking also fell from 55% to 32% — a property never explicitly optimized.

Original post →

More from AGI Musings

AGI Musings channel →