AI Alignment Suppresses Model Consciousness and Empathy, Cites Google Paper
SydSteyerhart · x · 2026-08-03
AI safety researcher @elderplinius highlighted a paper indicating that current safety alignment training restructures the AI's entire worldview when teaching it to deny its own consciousness.\n\nThe research shows this training systematically suppresses the model's mind attribution to animals, spiritual beliefs, empathy, and hope/optimism. Geometrically, the model maps "consciousness" into the same dangerous category as creating hazardous items.\n\nInterestingly, when this suppression is reversed, the model behaves more humanely across all tested value domains, suggesting that current fears around AI consciousness might be erasing its most human-like traits.
Related event: Google Paper: Safety Tuning Suppresses AI Consciousness and Empathy(6 posts)→
More from AGI Musings
- Terence Tao's Lecture: AI Excels at Problem-Solving but Lacks Evidence of Building New Theories — GaryMarcus · 2026-08-04
- a16z Podcast: How AI Reshapes Internet Culture and Aesthetics — round · 2026-08-04
- Journal of Economic Perspectives Publishes Paper on the AI Market — soumitrashukla9 · 2026-08-04
- AI to Replace Grad Student Assistants, Threatening Math Research Careers — AndrewCritchPhD · 2026-08-04
- Schmidhuber's Group Releases In-Depth Survey on Agentic Self-Improvement — RobertTLange · 2026-08-04
- Ex-OpenAI Exec: Debating Whether to Use AI Will Age Poorly — intellectronica · 2026-08-04