Debate erupts over Microsoft's 'dangerous' stance on model self-presentation
NinaPanickssery · x · 2026-09-18
After rgblong argued that Microsoft's takes on model self-presentation seem counterproductive and dangerous from a safety standpoint, Nina Panickssery pushed back: isn't this basically Roko's basilisk reasoning, and do you really have so little faith in aligning models to obedience that no lab should even try? The exchange continued into rgblong's specific concerns about embedding contradictory views of consciousness, goals, and self in models.
Related event: Microsoft's 'Model Welfare' Push Sparks AI Safety Debate(6 posts)→
More from AGI Musings
- Plinz: No one will pause AI — the real goal is keeping capable models away from the public — nptacek · 2026-09-18
- Debate: Should future people be discounted? Longtermism vs temporal discounting on X — NathanpmYoung · 2026-09-18
- "Parents liable for children" signs in Austria spark call to hold AI labs accountable — cnakazawa · 2026-09-18
- Anthropic says 80% of Claude's own code is now written by Claude — chrismattmann · 2026-09-18
- ArXiv 'proof' posts with undefined basics flood timelines, researcher warns of math slop — _onionesque · 2026-09-18
- 'Isn't managing the job?' — pushback on managers panicking over AI slop — beglen · 2026-09-18