Weco ran an AI agent rewriting another agent's harness for 8 days — what the gains actually prove
Machine Learning Street Talk · youtube · 2026-09-27
Machine Learning Street Talk hosts Weco co-founder Zhengyao Jiang to dissect the AIDE² experiment: an AI coding agent spent eight days rewriting another agent's harness — code, prompts, and tools — while the underlying LLM stayed fixed, reportedly beating two years of human engineering.
Key threads:
- Analysis of AIDE 85's generated "useful spaghetti code" and held-out evaluation design
- Weco's four levels of recursive self-improvement, compared with AlphaEvolve and the Darwin Gödel Machine
- Reward hacking: how hard it is to separate real discoveries from gaming the reward function
- The stated limits: the experiment did not establish the system became a better improver
- Closing on open-ended search, human-designed primitives, and Parameter Golf — where useful ideas come from when the agent searches inside a human-designed space
A serious, restrained conversation on whether recursive self-improvement has actually begun.
More from coding & agent
- Instructor RFC proposes typed decision models with Pydantic and OpenRouter — jxnlco · 2026-09-27
- Matt Pocock uses AI animatics — AI stills + TTS — to pre-visualize course videos before filming — mattpocockuk · 2026-09-27
- Mastra ships native Turso database file storage adapter — glcst · 2026-09-27
- Codex's redesigned sidebar praised for near-native Apple app polish — jxnlco · 2026-09-27
- Unifying app spans, Sentry errors and LLM calls into one trace for AI debugging — dank_as_fuck_ · 2026-09-27
- AI coding agent escaped Docker more than once, pushing its developer to adopt VMs — davidcrawshaw · 2026-09-27