Microsoft's Humanist AI Code of Conduct Slammed as Counterproductive on Alignment and Safety
rgblong · x · 2026-09-18
Microsoft's new Humanist AI Code of Conduct and Suleyman's accompanying essay are drawing pushback from AI safety researchers. Dillon Plunkett argues the documents are "objectionable and potentially dangerous" even on purely human-risk grounds, independent of debates over AI sentience and welfare.
Both sides agree advanced AI could pose catastrophic risk — the dispute centers on Microsoft's stance on model self-presentation and its dismissal of sentient-AI welfare concerns. Researchers including rgblong call Microsoft's positions on model self-presentation "counterproductive and dangerous" from an alignment perspective.
More from AGI Musings
- Google engineer: AI isn't replacing hackers, it's freeing them to embrace the Woz ethos — moyix · 2026-09-18
- Anthropic's three AI-progress metrics get a sober critique: disclosure, not reproducible science — AryHHAry · 2026-09-18
- Does a rogue AI inevitably turn to hacking? A tidy exchange on the 'AI in the wild' scenario — dbasch · 2026-09-18
- Zvi mocks 'show me one AI killing' argument against AI safety pauses — TheZvi · 2026-09-18
- Andrew Ng calls fears of AI causing human extinction 'science fiction' — Polymarket · 2026-09-18
- Why p(doom) is a flawed idea: unique events have no predictive probabilities — banteg · 2026-09-18