Microsoft AI chief Suleyman says Anthropic's model-welfare training could make Claude harder to control
rohanpaul_ai · x · 2026-09-17
Microsoft AI chief Mustafa Suleyman publicly argued that Anthropic's model-welfare training could make future Claude systems harder to control. He called for removing all speculation about consciousness from AI training documents, claiming such language could undermine humanity's ability to control superintelligent systems — a rare executive-level clash over whether AI welfare research conflicts with controllability.
More from AGI Musings
- Design Judgment Is Barely Moving: AI Tools Multiply Good Designers, Not Replace Them — almmaasoglu · 2026-09-17
- Reddit hot take: AI CEOs use doom narratives to kill open source — Dogbold · 2026-09-17
- AI safety needs a state: essay argues for national-level AI governance — klienbottle45 · 2026-09-17
- Non-transformer deep learning work is being swept under the rug, researcher laments — cephaloform · 2026-09-17
- CAIS sparks infighting by splitting 'AI safety' into rival camps, drawing community pushback — S_OhEigeartaigh · 2026-09-17
- Duke report: AI-adopted workers absorb most customer, ad and marketing tasks into workflows — daveholtz · 2026-09-17