Steering recursive self-improvement is 'close to nonsensical' amid alignment uncertainty
davidmanheim · x · 2026-09-07
AI safety researcher davidmanheim argues that suggestions to intervene and steer recursive self-improvement (RSI) are close to nonsensical: we still don't know what alignment in general would require or how to accomplish it, making RSI steering premature.
More from AGI Musings
- Stanford's Chris Potts revives 8-year-old "Deep RL Doesn't Work Yet" to counter AI skeptics — ChrisGPotts · 2026-09-07
- Physicist Sabine Hossenfelder: AI Has Eaten Mathematics, Physics Is Next — skdh · 2026-09-07
- "AI psychosis" is a lazy slur, argues researcher: Hinton, Chalmers and Anthropic take machine consciousness seriously — repligate · 2026-09-07
- Rumor: Anthropic's Claude solved Navier-Stokes; but is AI solving open problems just reward hacking? — burny_tech · 2026-09-07
- COT Backrooms: The Year After a Machine Solved Protein Folding, Told by the Scientists It Displaced — repligate · 2026-09-07
- AGI Definition Fight: Is GPT-6 Astra AGI, or Do People Just Want a Servant? — teortaxesTex · 2026-09-07