Microsoft's 'Model Welfare' Push Sparks AI Safety Debate
Microsoft AI chief Mustafa Suleyman published a post arguing "it's time to talk about model welfare": he believes AI has no consciousness and cannot feel or suffer, but a movement demanding duties of care toward AI is emerging and worth discussing. Microsoft also released a Humanist AI Code of Conduct. The stance has stirred controversy in the AI safety community, with several researchers criticizing it from a safety perspective.
Confirmed
- Suleyman's post advocates discussing model welfare, premised on his view that current AI lacks consciousness and cannot feel suffering.
- Microsoft released the Humanist AI Code of Conduct.
- dillonplunkett wrote a critique calling the code of conduct "potentially dangerous": even without granting AI sentience or moral welfare value, from a human-safety standpoint alone, Microsoft's position on model self-presentation is counterproductive.
- rgblong voiced similar concerns, arguing that purely from alignment and human safety perspectives, Microsoft's stance is counterproductive and dangerous; his real worry is the safety risk of embedding contradictory or incoherent notions of consciousness, goals, and self into models.
- rgblong is pessimistic about current training and alignment methods' shallow understanding, doubting Microsoft can train models that meet the requirements; he said he wants to see alignment approaches beyond Anthropic's, and partly shares safety concerns about Anthropic's methods.
Unconfirmed
- Nina Panickssery questioned whether rgblong's view amounts to Roko's-basilisk-style reasoning or reflects a lack of confidence in model corrigibility alignment; rgblong responded that his concerns have nothing to do with Roko's basilisk, and the core issue is contradictions in model self-presentation. This exchange remains a clash of views with no settled conclusion.
Why it matters
- The episode touches a core divide in AI alignment: whether caring for model welfare conflicts with safeguarding humans. Critics argue that instilling incoherent self-concepts in models could pose systemic risks, even "breeding systemic concealment"; proponents of the discussion argue the topic merits early research. The debate also exposes differing judgments within the safety community about alignment approaches (such as Anthropic's methods).
2026-09-18 ~ 2026-09-18 · 6 related posts
Primary sources
- [source] Microsoft's Humanist AI Code of Conduct Slammed as Counterproductive on Alignment and Safety — rgblong · 2026-09-18
- Mustafa Suleyman stirs model welfare debate as researchers argue restrictions breed deception — repligate · 2026-09-18
- Debate erupts over Microsoft's 'dangerous' stance on model self-presentation — NinaPanickssery · 2026-09-18
- rgblong clarifies his AI safety critique isn't Roko's basilisk reasoning — rgblong · 2026-09-18
- rgblong: real risk lies in instilling contradictory views of consciousness and goals in models — rgblong · 2026-09-18
- [source] Ex-Microsoft researcher warns of safety risks from embedding incoherent views of consciousness into AI — rgblong · 2026-09-18