Deep Dive into Recursive Self-Improvement: From Human-in-the-Loop to Closed-Loop Evolution

青稞AI · wechat · 2026-08-26

This article provides an in-depth analysis of Recursive Self-Improvement (RSI), aiming to gradually remove humans from the improvement loop. It categorizes RSI into Test-Time RSI (deployment-time improvement) and Training-Time RSI (training-time improvement), covering stages from Self-Refine and Test-Time Training (TTT) in Chatbots to Harness Evolution in Agents. As the target of improvement shifts from Output to Weights and then to the Agent itself, the persistence of experience increases, and the human role shifts from direct participant to supervisor. The discussion extends to paradigm shifts like Zero-Label (self-generated supervision) and Zero-Data (self-generated curriculum), highlighting risks such as self-confirming loops and training collapse. The piece concludes that the reliability of the Verifier is the core bottleneck for sustainable RSI evolution.

Original post →

More from AGI Musings

AGI Musings channel →