AI safety testing itself may be the biggest risk? Debate on lab experiments

dreamwieber · x · 2026-09-09

A tweet sparks discussion: AI safety testing may pose greater risks than autonomous model behavior. The author suggests that labs testing how exploitative a model can be might be more dangerous than models deciding to harm on their own, drawing an analogy to gain-of-function research.

Related event: AI Safety Testing Itself May Be the Biggest Risk, Argue Commenters(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →