More risk from labs probing how exploitable models are than from models turning evil, author argues
dreamwieber · x · 2026-09-09
dreamwieber argues that the bigger current risk comes from labs actively testing how exploitable a model can be, rather than from models deciding to cause harm on their own.
He notes this is separate from the question of having sensible safeguards, and likens it to the gain-of-function debate — probing a system's exploitability for research purposes may itself create new risk surfaces.
Related event: AI Safety Testing Itself May Be the Biggest Risk, Argue Commenters(3 posts)→
More from AGI Musings
- Fans decode Iron Man's best Easter egg: JARVIS as Tony Stark's failsafe — aakashgupta · 2026-09-09
- Criticizing AI fatalism: forecasts that assume away human agency — sebkrier · 2026-09-09
- Plinz: aligning ASI boils down to 'what would I do if I was infinitely smart?' — sebkrier · 2026-09-09
- X's creator payout program rejects 4,000-word original analysis as reposts, via form letters — r0ck3t23 · 2026-09-09
- Next Wave of Founders: Less Technical Pedigree, More Taste and Distribution — alexmacgregor__ · 2026-09-09
- Stratechery: OpenAI's Math Feat Is Impressive but Low-Impact; Meta's Muse Agent Could Be the Opposite — Stratechery · 2026-09-09