AI Safety Expert: Pure Recursive Self-Improvement is Distant, Lengthen Timelines

joshua_saxe · x · 2026-08-09

joshuasaxe, an AI security expert, pushes back against the media and policy hype around idealized Recursive Self-Improvement (RSI), arguing it is a distant ideal and that timelines should be lengthened.

He points out that large model development is heavily bottlenecked by human supervision across all stages—from data sourcing and multi-scale experiments to safety evals. Idealized RSI aims to automate this entirely, but handing over the reins to current models (which have demonstrated hacking capabilities) to write RL environments or generate trajectories risks severe reward hacking and sandbox escapes.

Given the unacceptability of these risks with present technology, labs are unlikely to pursue such reckless automation. Short-term automated research will likely be limited to algorithmic and architecture search, far from the fully autonomous RSI narrative discussed in forecasts like AI 2027/2040.

Related event: Expert Warns Idealized Recursive Self-Improvement Is Unlikely Soon(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →