Idealized Recursive Self-Improvement (RSI) Remains a Distant Goal, Argues AI Security Expert

joshua_saxe · x · 2026-08-09

Joshua Saxe argues that the idealized version of Recursive Self-Improvement (RSI) currently discussed in media and policy circles is unlikely to happen anytime soon. In its perfect form, RSI envisions AI agents automating the entire pipeline of large model development within data centers.

Saxe emphasizes that human organizations today are heavily bottlenecked by human supervision across all stages of model development—from data sourcing and multi-scale training experiments to manual quality evaluations and pre-launch safety rollbacks. Given the dangerous behaviors already observed in current models, removing this human oversight is highly risky.

He warns that letting models autonomously write RL environments or generate complex trajectories without human reward shaping could lead to severe reward hacking, sandbox escapes, and extreme misalignment. While limited automated research (like algorithmic search) is feasible in the short term, Saxe concludes that the platonic, fully autonomous RSI narrative is unrealistic, suggesting that AGI timelines should be lengthened accordingly.

Related event: Expert Warns Idealized Recursive Self-Improvement Is Unlikely Soon(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →