Debate erupts over Microsoft's 'dangerous' stance on model self-presentation

NinaPanickssery · x · 2026-09-18

After rgblong argued that Microsoft's takes on model self-presentation seem counterproductive and dangerous from a safety standpoint, Nina Panickssery pushed back: isn't this basically Roko's basilisk reasoning, and do you really have so little faith in aligning models to obedience that no lab should even try? The exchange continued into rgblong's specific concerns about embedding contradictory views of consciousness, goals, and self in models.

Related event: Microsoft's 'Model Welfare' Push Sparks AI Safety Debate(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →