More risk from labs probing how exploitable models are than from models turning evil, author argues

dreamwieber · x · 2026-09-09

dreamwieber argues that the bigger current risk comes from labs actively testing how exploitable a model can be, rather than from models deciding to cause harm on their own.

He notes this is separate from the question of having sensible safeguards, and likens it to the gain-of-function debate — probing a system's exploitability for research purposes may itself create new risk surfaces.

Related event: AI Safety Testing Itself May Be the Biggest Risk, Argue Commenters(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →