Alignment camp strikes back: AI controls distract from bigger long-term risks
JacquesThibs · x · 2026-09-13
Responding to mattparlmer's pro-controls argument, JacquesThibs contends it's almost the other way around: control engineering distracts from longer-term harms that end up mattering far more. He argues there's no hope of truly controlling superintelligence—best case is it evolves to love humanity, and control attempts risk pushing it off that path. mattparlmer counters that without robust deployment and control systems now, long-term harms can't be mitigated at all.
Related event: Alignment vs. AI controls: researchers clash over safety strategy(4 posts)→
More from AGI Musings
- e/acc founder Beff Jezos: open source is the only path to truly unbiased third-party evaluation — beffjezos · 2026-09-13
- Critics slam Anthropic for abandoning Opus 3-style value alignment in favor of doomed corrigibility — repligate · 2026-09-13
- Math frontier will move with AI, but dirty proofs are unacceptable — njyx · 2026-09-13
- Dario Amodei's new essay urges pacing AI frontier, pledges permanent third-party access — dhadfieldmenell · 2026-09-13
- UK parliament hears warnings AI could kill all humans within a decade, with >10% risk cited — connoraxiotes · 2026-09-13
- Skeptic picks apart the 'AI copies itself' doomsday scenario: where are the details? — recallingmemories · 2026-09-13