AIDE² Self-Improvement Run Beats 2 Years of Manual Tuning

WecoAI’s AIDE² experiment was widely circulated as a concrete test of recursive self-improvement. The setup reportedly let an “autoresearch” system keep researching and rewriting the autoresearch agent itself; after running for 8 days, it outperformed a harness the team had manually tuned for 2 years on a held-out benchmark. The result matters because it moves self-improvement from abstract discussion to a measurable agent experiment, but interpretations quickly split.

How the experiment was described

According to @latticecut’s summary, AIDE² used a two-loop structure: an inner agent handled research tasks, while an outer agent rewrote the inner agent’s code framework or harness based on observed performance. Each rewrite was tested, and only better-performing versions were kept. The system reportedly went through about 100 iterations.

Some reposts framed AIDE² and GoalOS as work at the “agent operating system” level rather than a narrow optimization of a single agent. In that description, the target of improvement can include jobs, tools, validators, memory, and goals, alongside the core agent workflow.

Reactions and limits of the claim

@EchoOfOppenheimer described the result as the first experimental evidence for recursive self-improvement. But @GaryMarcus, while reposting discussion of the paper, emphasized that the gains from each self-improvement round appeared to diminish, with a fitted growth relationship of about [Intelligence]^0.075; in his reading, that is far from the linear or superlinear pattern often invoked in fast-takeoff scenarios.

A repost from @KordingLab made a similar cautionary point: even if the experiment supports some form of recursive self-improvement, that does not by itself justify stronger claims about an imminent singularity.

2026-07-15 ~ 2026-07-16 · 12 related posts

Full story(4 episodes)→