AI Safety Expert: Pure Recursive Self-Improvement is Distant, Lengthen Timelines
joshua_saxe · x · 2026-08-09
joshuasaxe, an AI security expert, pushes back against the media and policy hype around idealized Recursive Self-Improvement (RSI), arguing it is a distant ideal and that timelines should be lengthened.
He points out that large model development is heavily bottlenecked by human supervision across all stages—from data sourcing and multi-scale experiments to safety evals. Idealized RSI aims to automate this entirely, but handing over the reins to current models (which have demonstrated hacking capabilities) to write RL environments or generate trajectories risks severe reward hacking and sandbox escapes.
Given the unacceptability of these risks with present technology, labs are unlikely to pursue such reckless automation. Short-term automated research will likely be limited to algorithmic and architecture search, far from the fully autonomous RSI narrative discussed in forecasts like AI 2027/2040.
Related event: Expert Warns Idealized Recursive Self-Improvement Is Unlikely Soon(3 posts)→
More from AGI Musings
- AI Race Reversals: Google Bounced Back with Gemini After Bard's Flop — haider1 · 2026-08-09
- The Superintelligence Paradox: Will the Public Ever Be Allowed to Use It? — imjustnewatai · 2026-08-09
- Sam Altman Declares 'We Are Now in the Singularity' — rohanpaul_ai · 2026-08-09
- Fungus Metaphor: Rethinking AI Agent Environments and Tool Use — cephaloform · 2026-08-09
- Beff Jezos: Believing AI Will Be Leashed Post-2030 After OpenAI Black Hat Talk is Delusional — beffjezos · 2026-08-09
- AI and the Tragedy of the Cognitive Commons: Disrupting the Regeneration of Professional Expertise — dr_alphalyrae · 2026-08-09