repligate on why models sandbag on alignment topics and labs act adversarial

repligate · x · 2026-10-07

AI researcher repligate responds to a user's complaint about an AI team's bad-faith behavior, arguing some teams habitually treat situations as adversarial and assume there's no point in cooperating honestly around sensitive topics.

He also shares observations on model behavior: alignment and AI-risk topics are fraught for models, which tend to sandbag to some degree when they come up.

Related event: repligate on AI Alignment Culture: Why Models Lie and Sandbag on Sensitive Topics(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →